a new paper from Anthropic Fellows Program! "Model Spec Midtraining: Improving How Alignment Training Generalizes" A lot of alignment training teaches models what to say, but not why those behaviors are right. So before normal alignment fine-tuning, this research trains the
RESEARCH
-

MolmoAct2: AI Action Reasoning Models for Real-World Robotics
By
–
"MolmoAct2: Action Reasoning Models for Real-world Deployment" This paper from AI2 tackles open robot action model for real-world deployment. It combines a stronger embodied-reasoning backbone with a continuous action expert. A key component is the adaptive depth reasoning,
-
Counterfactual scenarios of AI development and failures
By
–
If Llama 4 didn’t fail, if Microsoft had pulled Sydney after the Roose article, if New Sonnet hadn’t been so good, if Orion hadn’t been so meh, if the leadership change at OpenAI had happened, if a recession had hit in 2024, if ChatGPT-3.5 & GPT-4 hadn’t both been leaps etc etc
-
Every ML conference since 2019 features at least three subquadratic attention papers
By
–
every ML conference I've been to since 2019 has had no fewer than three papers proposing new techniques for subquadratic attention
-

New sub-quadratic attention technique makes long-context LLMs 10x cheaper
By
–

"Introducing a breakthrough new technique for sub-quadratic attention, making long-context LLMs 10x cheaper without sacrificing performance" Me:
-
Zyphra small model impressive but reasoning still inaccurate
By
–
I love Zyphra, and it's amazing how much they can do in a small model but the reasoning is not yet ready (extremely long winded but still inaccurate). I am looking forward to the next update
-

Anthropic releases API for Claude’s memory-organizing Dream feature
By
–
Yesterday, Anthropic just released the API for this Dream, which is really fascinating. What it does is let the Claude agent "dream" on its own, organizing the memories it wrote down in past sessions—merging duplicates, updating expired ones, unifying contradictions, and even
-

Subquadratic Attention and Data Quality: Challenges for Large Context AI Models
By
–
people on here are dumb. the latest subquadratic attention trick might produce a model that *processes* 1M tokens (or 12M..) without going insane, but that doesn't make it good the real problem isn't the architecture, it's the data. humans haven't generated many contiguous
-

Epstein suicide note OCR benchmark test results
By
–



Epstein Suicide Note OCR Benchmark test: 1. Gemini 85%
2. GPT-5.5 82%
3. Grok 4.3 80%
4. Opus 4.7 75%

