8. Synthetic Dataset Generation for RAG Evaluation with Multi-Agent Systems The paper proposes a modular, three-agent pipeline that auto-generates synthetic QA datasets for evaluating RAG systems while enforcing privacy.
@dair_ai
-

Parallel Graph-Retrieval-Augmented Reasoning for Medical Knowledge
By
–
5. Parallel Graph-Retrieval-Augmented Reasoning A test-time reasoning framework that replaces a single linear chain with multiple parallel, entity-grounded chains over medical knowledge graphs.
-

Evaluating Language Models on Real Unsolved Questions
By
–
7. Assessing Language Models on Unsolved Questions The paper introduces a new evaluation paradigm that tests models on real unsolved questions from the wild, rather than on fixed-answer exams.
-

Memory-R1: Framework Teaching LLM Agents Memory Management
By
–
6. Memory-R1 A framework that teaches LLM agents to decide what to remember and how to use it.
-

Jet-Nemotron: Hybrid LLM Architecture with Optimized Attention
By
–
4. Jet-Nemotron A hybrid-architecture LM family: starting from a frozen full-attention model, the authors search for where to keep full attention, which linear-attention block to use, and which hyperparameters match hardware limits.
-

Memory-based LLM Agent Fine-tuning Without Weight Updates
By
–
3. Fine-tuning LLM Agents without Fine-tuning LLMs A memory‑based learning framework that lets deep‑research agents adapt online without updating model weights.
-
Deep Think with Confidence: Pruning Weak Reasoning Paths
By
–
2. Deep Think with Confidence
— DAIR.AI (@dair_ai) 31 août 2025
A lightweight test-time method that uses model-intrinsic confidence to prune weak reasoning paths, improving both accuracy and token efficiency for self-consistency ensembles.https://t.co/X9M1n0mKok2. Deep Think with Confidence A lightweight test-time method that uses model-intrinsic confidence to prune weak reasoning paths, improving both accuracy and token efficiency for self-consistency ensembles.
-

Reinforcement Learning Techniques for LLM Reasoning Evaluated
By
–
10. A Deep Dive into RL for LLM Reasoning This paper reviews and rigorously re-evaluates reinforcement learning techniques for LLM reasoning, addressing inconsistencies caused by varied setups and unclear guidelines.
-

Efficient LLM Architectures Beyond Traditional Transformers Survey
By
–
9. A Survey on Efficient Architectures for LLMs Reviews advances in efficient LLM architectures beyond traditional transformers, including linear and sparse sequence models, efficient attention variants, sparse MoEs, hybrid designs, and diffusion-based.
-

GLM-4.5: Open Mixture-of-Experts for Agentic AI
By
–
8. GLM-4.5 An open Mixture‑of‑Experts family that targets a single model excelling across agentic tool use, complex reasoning, and real‑world coding.
