Autonomous Memory Management in LLM Agents LLM agents struggle with long-horizon tasks due to context bloat. As interaction history grows, computational costs explode, latency increases, and reasoning degrades from distraction by irrelevant past errors. The standard approach
@dair_ai
-

Meta Lab’s Self-Evolving AI Agents Show Complex Reasoning Emergence
By
–
Super interesting paper from Meta Superintelligence Labs. This work suggests complex reasoning and search capabilities can emerge solely through self-evolution, challenging the assumption that human supervision is necessary for advanced agent abilities. Let's break down the
-

Improving Self-Evolving LLM Agents Learning Capabilities
By
–
On building more powerful self-evolving agents. LLM agents struggle to learn from experience after deployment. Fine-tuning is expensive and causes catastrophic forgetting. RAG retrieves based on semantic similarity alone, often pulling noise instead of what actually works.
-

Agent-as-a-Judge: Robust LLM Evaluation Framework
By
–
A Survey on Agent-as-a-Judge Great report to learn about agentic judges used for planning, tool-augmented verification, multi-agent collaboration, and persistent memory to enable more robust, verifiable, and nuanced evaluations. With careful crafting, it's possible to build LLM
-

Deep Delta Learning: Novel Residual Connection Operator
By
–
10. Deep Delta Learning Deep Delta Learning introduces a novel “Delta Operator” that generalizes residual connections by modulating the identity shortcut with a learnable, data-dependent geometric transformation.
-

SWE-EVO Benchmark for Long-Horizon Software Evolution Tasks
By
–
9. SWE-EVO SWE-EVO introduces a benchmark for evaluating coding agents on long-horizon software evolution tasks that require multi-step modifications spanning an average of 21 files per task.
-

SciSciGPT: Open-Source AI Research Collaborator for Science
By
–
8. SciSciGPT SciSciGPT is an open-source AI collaborator that uses the science of science domain as a testbed for LLM-powered research tools.
-

Confucius Code Agent: Large-Scale Codebase Engineering
By
–
7. Confucius Code Agent Confucius Code Agent (CCA) is a software engineering agent designed to operate on large-scale codebases.
-

Nemotron-Cascade: Cascaded Reinforcement Learning for Reasoning Models
By
–
4. Nemotron-Cascade Nemotron-Cascade introduces cascaded domain-wise reinforcement learning (Cascade RL) to build general-purpose reasoning models capable of operating in both instruct and deep thinking modes.
-
LLMs Evolve Competing Programs in Digital Red Queen Algorithm
By
–
3. Adversarial Program Evolution with LLMs
— DAIR.AI (@dair_ai) 11 janvier 2026
Digital Red Queen (DRQ) introduces an algorithm where LLMs evolve assembly-like programs called “warriors” that compete for control of a virtual machine in the game of Core War.https://t.co/u2OBC5HFxV3. Adversarial Program Evolution with LLMs Digital Red Queen (DRQ) introduces an algorithm where LLMs evolve assembly-like programs called “warriors” that compete for control of a virtual machine in the game of Core War.
