Great paper on long-term memory for LLM agents. (bookmark it) Coarse summaries drift and unconstrained updates corrupt, so AtomMem makes the unit of memory small. A Fact Executor pulls high-value atomic facts out of long interactions, organizes them into hierarchical event
RESEARCH
-

DexJoCo: New Benchmark for Task-Oriented Dexterous Manipulation
By
–
Can robotic hands truly handle complex tasks like tool use and bimanual coordination? CASIA, SJTU, MBZUAI, PKU, and CUHK introduce DexJoCo — a new benchmark and toolkit built for task-oriented dexterous manipulation. It features 11 human-grounded tasks, a low-cost data
-
Karpathy’s RL prediction: reward functions unreliable, need knowledge-guided review
By
–
Karpathy's prediction about RL is coming true now!
— Akshay 🚀 (@akshay_pachaar) 19 juin 2026
He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a… https://t.co/0REApdfBUG pic.twitter.com/uAfW9yvn3mKarpathy's prediction about RL is coming true now! He called reward functions unreliable and argued that a single reward number is too low-dimensional to teach an agent what "good" means for complex tasks. To solve this, Agents need a knowledge-guided review as a
-
Government shuts down Claude AI that beat Opus benchmarks
By
–
7 days ago a government order took the most powerful Claude offline.
It beat Opus on almost every benchmark.
It did a 2-month job for Stripe in a day.
Now nobody can touch it. I still run its best move on Opus 4.8 every day.
And tell me: should one government get to switch off -
Jagged intelligence makes predicting model behavior harder
By
–
Jagged intelligence in one tweet haha. It's harder and harder to predict what will work and won't the more we grow the harnesses around these models.
-

600M model outperforms 397B and Sonnet 4.5
By
–

600M that beats a 397B and Sonnet 4.5 Small and specialized models FTW
-

GLM-5.2 performs poorly on Bullshit Benchmark, like other Series 5 models
By
–
GLM-5.2 did not perform very well on the Bullshit Benchmark – a level similar to that of their other Series 5* models.
-

The Foundations of Transformers, Pretraining, Distillation, and World Models in 1991
By
–
In 1991, the foundations of Transformers, pretraining, distillation, and world models were already being built. These have helped shape my own thinking, from my time at Google Brain to our work on Improvement.
-

Rapid AI evolution: Year-old tools face new specialized competitors
By
–
The AI landscape is moving incredibly fast. Tools that were considered cutting-edge just a year ago are already facing new competitors focused on specific use cases like research, presentations, note-taking, content creation, and workflow automation. What's interesting is that
