Sure, reasoning models can learn when to think, but how? A team from Tsinghua University find that NoThinking is a better choice for relatively simple tasks in terms of both performance and efficiency. Nothing surprise here. but they propose AdaptThink too! AdaptThink is a
LLMS
-

Fractured Chain-of-Thought Reasoning Advances AI Reasoning
By
–
Fractured Chain-of-Thought Reasoning
Paper: https://
arxiv.org/pdf/2505.12992 -

Fractured Sampling Boosts Reasoning Inference Without Retraining
By
–
Full and Long CoT boost reasoning by expanding intermediate steps—but at a high token cost. Not ideal for latency- or cost-sensitive apps. This paper introduces Fractured Sampling, a practical inference-time technique that turbocharges reasoning without retraining, using far
-

Qwen3-32B Performance Benchmark Against Leading AI Models
By
–
Qwen3-32B vs. leading models eval https://
artificialanalysis.ai/?models=gpt-4-
1o3llama-4-scoutllama-4-maverickgemma-3-27bgemini-2-5-proclaude-3-7-sonnet-thinkingclaude-3-7-sonnetmistral-medium-3deepseek-r1deepseek-v3-0324grok-3grok-3-mini-reasoningqwen3-32b-instruct-reasoning
… -

Cerebras Qwen3 achieves 99% latency reduction versus o3
By
–
Artificial Analysis measured the time to first token of every reasoning model from o3 to R1. DeepSeek R1 = 103 sec
Qwen3-32B on Cerebras = 1.1 sec We give you R1 level intelligence with 99% latency reduction. Try it: https://
inference.cerebras.ai -
Compare Qwen, Llama, DeepSeek Models on Cerebras Platform
By
–
Choose a model (
@Alibaba_Qwen 3 32B, @AIatMeta Llama 3.3 70B, Llama 4 Scout, @deepseek_ai R1 Distill Llama 70B) – https://
poe.com/search?q=cereb
ras
… Add a prompt template Connect your data Chain it with other tools -
Gemma 3n: Multimodal AI Model Running on 2GB RAM
By
–
Meet Gemma 3n, a model that runs on as little as 2GB of RAM 🤯 It shares the same architecture as Gemini Nano, and is engineered for incredible performance. We added audio understanding, so now it’s multimodal, fast and lean, and runs on-device (no cloud connection required!) pic.twitter.com/2FyzJHVGZa
— Google AI (@GoogleAI) 20 mai 2025Meet Gemma 3n, a model that runs on as little as 2GB of RAM It shares the same architecture as Gemini Nano, and is engineered for incredible performance. We added audio understanding, so now it’s multimodal, fast and lean, and runs on-device (no cloud connection required!)
-
Google Gemini 2.5 Scores 49% on USA Math Olympiad
By
–
A few months ago, the best LLM scored 5% on the USA Math Olympiad test. Models have been rapidly improving. Today, Google Gemini 2.5 scored 49%, which is better than 75% of the people who took the test (roughly the top 250 students in the USA).
-

LLMs Surpass Benchmarks: Converting AI Capabilities Into Business Value
By
–
LLMs are blowing through benchmarks faster and faster. Next up, converting capabilities into business value.
-
Testing AI Model Performance with Extended Prompt Engineering
By
–
I'll try prompting with some really long verse and see how it handles
