5). Rethinking Reflection in Pre-Training The authors introduce adversarial reasoning tasks to show that self-reflection and correction capabilities steadily improve as compute increases, even in the absence of supervised post-training.
@dair_ai
-

New RL Training Strategy Enhances LLM Reasoning Conciseness
By
–
4). Concise Reasoning via RL This new paper proposes a new training strategy that promotes concise and accurate reasoning in LLMs using RL.
-

AI Scientist V2 Autonomously Generates Workshop-Accepted Research
By
–
1). The AI Scientist V2 The AI Scientist-v2 refines and extends its predecessor to achieve a new milestone: autonomously generating a workshop-accepted research manuscript.
-
Top AI Research Papers Week April 7-13 2026
By
–
Here are the top AI Papers of the Week (April 7 – 13): – NoProp
– The AI Scientist V2
– Concise Reasoning via RL
– Rethinking Reflection in Pre-Training
– Efficient KG Reasoning for Small LLMs
– Agentic Knowledgeable Self-awareness (bookmark for later) Read on for more: -

Efficient Reasoning in LLMs: Balancing Performance and Cost
By
–
9). A Survey of Efficient Reasoning for LLMs This survey focuses on reasoning economy in LLMs, analyzing how to balance deep reasoning performance with computational cost.
-

Hidden Factual Knowledge Encoded in Large Language Models
By
–
10). Hidden Factual Knowledge in LLMs This study introduces a framework to measure hidden knowledge in LLMs, showing that models encode significantly more factual information internally than they express in outputs, up to 40% more.
-

Open Deep Search: Open-Source AI Search Framework Rivals GPT-4o
By
–
7). Open Deep Search Introduces Open Deep Search (ODS), an open-source search AI framework that rivals top proprietary systems like GPT-4o Search Preview and Perplexity Sonar.
-
Z1: Dynamic Reasoning Method Makes LLMs More Compute-Efficient
By
–
8). Z1 Z1 is a new method for making LLMs more compute-efficient at test time, especially during reasoning. It train LLMs with short and long code-based reasoning trajectories, and then dynamically adjusts reasoning depth during inference.
-

MedAgentSim: Automated LLM Hospital Simulation for Doctor-Patient Interactions
By
–
6). MedAgentSim MedAgentSim is a fully automated, open-source hospital simulation where LLM-powered agents simulate doctor-patient interactions in dynamic diagnostic settings.
-

Why LLMs Focus Attention on First Token
By
–
5). Why do LLMs Attend to First Token? This new paper explains why LLMs obsessively focus attention on the first token — a phenomenon known as an attention sink.
