6. Self-Challenging Language Model Agents Proposes a novel self-improvement method for multi-turn tool-use LLM agents, called the Self-Challenging Agent (SCA). It trains LLMs entirely from tasks they generate themselves, avoiding the need for human-annotated tasks or
@dair_ai
-

Evaluating LLM Knowledge and Reasoning with KI InfoGain
By
–
3. Knowledge or Reasoning Introduces a fine-grained evaluation framework to dissect LLM thinking into two components: knowledge correctness and reasoning informativeness, measured via Knowledge Index (KI) and Information Gain (InfoGain), respectively.
-

OpenThoughts3: SFT Data Curation for Open-Source Reasoning
By
–
4. Open Thoughts This paper presents OpenThoughts3, a systematic recipe for curating supervised fine-tuning (SFT) data that advances the performance of open-source reasoning models.
-
Information-Theoretic Framework for LLM Semantic Knowledge Organization
By
–
2. From Tokens to Thoughts This paper introduces an information-theoretic framework to examine whether LLMs organize semantic knowledge like humans, balancing compression and meaning.
-

GRIT: Teaching MLLMs Grounded Visual Reasoning with Images
By
–
10. Teaching MLLMs to Think with Images GRIT is a new method that enables MLLMs to perform grounded visual reasoning by interleaving natural language with bounding box references.
-

ARC-AGI-2: New Benchmark Pushes AI Reasoning Boundaries
By
–
9. ARC-AGI-2 ARC-AGI-2 is a new benchmark designed to push the boundaries of AI reasoning beyond the original ARC-AGI.
-

MedBrowseComp: LLM Agents Medical Fact-Finding Benchmark
By
–
8. MedBrowseComp MedBrowseComp is a new benchmark designed to evaluate LLM agents’ ability to perform complex, multi-hop medical fact-finding by browsing real-world, domain-specific web resources.
-
AdaptThink: RL Framework for Dynamic Reasoning Strategy Selection
By
–
7. AdaptThink This paper introduces AdaptThink, an RL framework designed to help reasoning models decide when to use detailed chain-of-thought reasoning (“Thinking”) versus directly producing an answer (“NoThinking”), based on task difficulty.
-

LLM Reasoning in Dynamic Environments Beyond Static Benchmarks
By
–
6. Towards a Deeper Understanding of Reasoning in LLMs This paper investigates whether LLMs can adapt and reason in dynamic environments, moving beyond static benchmarks.
-
AI Model Predicts Immunotherapy Treatment Response Outcomes
By
–
5. AI Predicts Immunotherapy Outcomes Across Cancers and Treatments Introduces COMPASS, a concept bottleneck-based foundation model that predicts patient response to immune checkpoint inhibitors (ICIs) using tumor transcriptomic data.