10). Inner Thinking Transformers – A new method that enhances reasoning efficiency in small-scale LLMs via dynamic depth scaling. ITT aims to mitigate parameter bottlenecks in LLMs, providing scalable reasoning efficiency without expanding model size.
@dair_ai
-

The Danger of Overthinking in Large Reasoning Models
By
–
9). The Danger of Overthinking – This paper investigates overthinking in Large Reasoning Models (LRMs)—a phenomenon where models prioritize extended internal reasoning over interacting with their environment.
-

Open-Reasoner-Zero: Efficient RL Framework Outperforms DeepSeek-R1
By
–
7). Open-Reasoner-Zero – an open-source large-scale minimalist RL framework to enhance reasoning. Achieves significant scalability requiring only 1/30th of the training steps of DeepSeek-R1-Zero-Qwen-32B to outperform it on GPQA Diamond.
-

MoBA: New Attention Mechanism for Efficient Long-Context LLMs
By
–
8). MoBA – A new attention mechanism that enhances efficiency in handling long-context sequences for LLMs while maintaining strong performance.
-

Text Data Augmentation Techniques for Large Language Models
By
–
10). Survey: Text Data Augmentation for LLMs This comprehensive survey covers text data augmentation techniques for LLMs.
-

Self-MoA: Single Model Ensembles Outperform Multi-Model Mixtures
By
–
7). Rethinking Mixture-of-Agents This paper asks: is mixing different LLMs actually helpful, or are we better off ensembling one top model’s outputs? The surprising answer: “Self-MoA” (single-model ensemble) often wins over multi-model ensembles.
-

MaAS: Universal Agentic Supernet for Dynamic Agent Teams
By
–
8). MaAS MaAS (Multi-agent Architecture Search) learns a universal “agentic supernet” from which it can spawn an optimal agent team on the fly for each query.
-

Enhancing Reasoning Capabilities in Large Language Models
By
–
9). Advancing Reasoning in LLMs This survey paper provides a timely overview of emerging methods to enhance reasoning capabilities in LLMs.
-

Long Chain-of-Thought Reasoning in LLMs: RL and Scaling
By
–
6). Demystifying Long Chain-of-Thought Reasoning in LLMs This work investigates how LLMs develop extended CoT reasoning, focusing on RL and compute scaling.
-
Top AI Papers of the Week: Reasoning and Agent Advances
By
–
Here are the top AI Papers of the Week (Feb 3-9): – s1
– OmniHuman-1
– Less Is More for Reasoning
– Advancing Reasoning in LLMs
– Rethinking Mixture-of-Agents
– Chain-of-Associated-Thoughts Read on for more:
