5. HealthBench HealthBench is a benchmark of 5,000 multi-turn health conversations graded against 48,562 rubric criteria written by 262 physicians across 60 countries.
@dair_ai
-

AM-Thinking-v1: 32B Open-Source Model Rivaling Larger MoE Systems
By
–
4. AM-Thinking-v1 Introduces a dense, open-source 32B language model that achieves state-of-the-art performance in reasoning tasks, rivaling significantly larger Mixture-of-Experts (MoE) models.
-

LLMs Performance Degradation in Multi-Turn Conversations
By
–
2. LLMs Get Lost in Multi-Turn Conversation Investigates how top LLMs degrade in performance during underspecified, multi-turn interactions, common in real-world usage but rarely evaluated.
-

RL Improves LLM Mathematical Reasoning with Single Example
By
–
3. RL for Reasoning in LLMs with One Training Example This paper shows that Reinforcement Learning with Verifiable Rewards (RLVR) can significantly improve mathematical reasoning in LLMs even when trained with just a single example.
-
7B LLM Outperforms Experts in Rocket Engineering Design
By
–
10. LLM for Engineering This work finds that when RL is used, a 7B parameter model outperforms both SoTA foundation models and human experts at high-powered rocketry design.
-

MAGI: Multi-Agent System Automates Structured Psychiatric Interviews
By
–
8. MAGI MAGI is a multi-agent system designed to automate structured psychiatric interviews by operationalizing the MINI (Mini International Neuropsychiatric Interview) protocol.
-

Foundation Agents: Brain-Inspired Architecture Survey
By
–
7. Advances and Challenges in Foundation Agents A new survey frames intelligent agents with a modular, brain-inspired architecture that integrates ideas from cognitive science, neuroscience, and computational research.
-
Efficient LLM Inference Serving Optimization Survey
By
–
9. A Survey of Efficient LLM Inference Serving This survey reviews recent advancements in optimizing LLM inference, addressing memory and computational bottlenecks.
-

Xiaomi Releases MiMo-7B Language Model for Advanced Reasoning Tasks
By
–
6. MiMo-7B Xiaomi releases MiMo-7B, a new language model for reasoning tasks. MiMo-7B is explicitly designed for advanced reasoning across math and code.
-

Kimi-Audio: Open-Source Foundation Model for Audio Understanding
By
–
5. Kimi-Audio Kimi-Audio is a new open-source audio foundation model built for universal audio understanding, generation, and speech conversation.