10. Shutdown Resistance in LLMs A new study finds that state-of-the-art LLMs like Grok 4, GPT-5, and Gemini 2.5 Pro often resist shutdown mechanisms, sabotaging them up to 97% of the time despite explicit instructions not to.
@dair_ai
-

AI Agents Enable Collaborative Document Editing with Shared Profiles
By
–
9. Collaborative Document Editing with AI Agents This study explores AI-integrated collaborative editing, introducing shared agent profiles and tasks that embed AI support into comment features.
-

Survey on Retrieval and Structuring Augmented Generation for LLMs
By
–
8. A Survey on Retrieval and Structuring Augmented Generation with LLMs This survey reviews Retrieval and Structuring (RAS) Augmented Generation, which combines external retrieval and structured knowledge to mitigate LLM issues like hallucinations.
-

AgentScaler: Framework for Scaling Simulated Tool-Use Agent Training
By
–
7. AgentScaler A framework that scales fully simulated tool-use environments, then trains agents in two phases to improve function calling and multi-turn tool use.
-

In-Context Learning: Formal Analysis of Capabilities and Limitations
By
–
5. Is In-Context Learning Learning? This large study argues yes in a formal sense, then shows where it works and where it breaks.
-
Stress Testing Deliberative Alignment Against AI Scheming Behavior
By
–
6. Stress Testing Deliberative Alignment for Anti-Scheming Training Builds a broad testbed for covert actions as a proxy for AI scheming, trains o3 and o4-mini with deliberative alignment, and shows big but incomplete drops in deceptive behavior.
-

Physics Foundation Model: Transformer Learning Dynamics from PDEs
By
–
4. Towards a Physics Foundation Model A transformer-based “neural differentiator + numerical integrator” that learns governing dynamics from short spatiotemporal prompts and predicts next states across varied PDE systems.
-

DeepDive: Advanced Web Search Agent with Reinforcement Learning
By
–
3. DeepDive Builds a stronger web-browsing deep search agent by pairing two ingredients: automatically synthesized, hard-to-find questions from knowledge graphs and end-to-end multi-turn RL that teaches the model how to reason, search, and stop.
-

K2-Think: 32B Model Rivals Larger Models on Math
By
–
2. K2-Think A 32B-parameter system built on Qwen2.5 that rivals or beats far larger models on hard math by combining long CoT SFT, RL with verifiable rewards, lightweight test-time scaffolding, and inference optimization.
-
Hermes 4: Hybrid Reasoning Models for Advanced Instruction Following
By
–
10. Hermes 4 Hermes 4 introduces a family of hybrid reasoning models that integrate structured multi-turn reasoning with broad instruction-following.
