9). Can LLMs Design Good Questions? – systematically evaluates the quality of questions generated with LLMs; here are the main findings: 1) there is a strong preference for asking about specific facts and figures in both LLaMA and GPT models…
@dair_ai
-

Process Reinforcement Implicit Rewards Framework Language Models
By
–
8). Process Reinforcement through Implicit Rewards – a framework for online reinforcement learning that uses process rewards to improve language model reasoning…
-

Cosmos World Foundation Model for Physical AI Training
By
–
7). Cosmos World Foundation Model – a framework for training Physical AI systems in digital environments before real-world deployment.
-

rStar-Math: Code-Augmented CoT for Enhanced Math Reasoning
By
–
6). rStar-Math – a new approach proposes three core components to enhance math reasoning: 1) a code-augmented CoT data synthesis method involving MCTS to generate step-by-step verified reasoning trajectories which are used to train the policy SLM…
-

Meta Chain-of-Thought Extends AI System 2 Reasoning Capabilities
By
–
5). Towards System 2 Reasoning – proposes Meta Chain-of-Thought (Meta-CoT), which extends traditional Chain-of-Thought (CoT) by modeling the underlying reasoning required to arrive at a particular CoT.
-

Agent Laboratory: LLM Agents Advancing Research Process
By
–
2). Agent Laboratory – an approach that leverages LLM agents capable of completing the entire research process; the main findings are: 1) agents driven by o1-preview resulted in the best research outcomes…
-

Long Context LLMs Outperform RAG in Question-Answering
By
–
3). Long Context vs. RAG for LLMs – performs a comprehensive evaluation of long context (LC) LLMs compared to RAG systems; the three main findings are: 1) LC generally outperforms RAG in question-answering benchmarks…
-

Search-o1 Framework Combines Reasoning Models with Agentic Search
By
–
4). Search-o1 – a framework that combines large reasoning models (LRMs) with agentic search and document refinement capabilities to tackle knowledge insufficiency…
-

DRT-o1 Applies Long-Chain Reasoning to Machine Translation
By
–
10). DRT-o1 – applies long chain-of-thought reasoning to machine translation, particularly for handling metaphors and similes across different cultures.
-

Reinforcement Learning Overview: Comprehensive Guide
By
–
9). Reinforcement Learning Overview – presents a comprehensive overview of reinforcement learning.
