We kept MRCR in the system card for scientific honesty, but we've actually been phasing it out slowly. Two reasons: (1) it's built around stacking distractors to trick the model, which isn't how people actually use long context, and (2) we care more about applied
RESEARCH
-

MRCR Phased Out: Focus Shifts to Applied Long-Context
By
–
We kept MRCR in the system card for scientific honesty, but we've actually been phasing it out slowly. Two reasons: (1) it's built around stacking distractors to trick the model, which isn't how people actually use long context, and (2) we care more about applied long-context
-
Microsoft AI Tools Design Insights for Industry Impact
By
–
Super proud of the work that went into this. These insights could have important implications for how we and the industry design these tools – making sure what we build is what people actually need. An important milestone for @MicrosoftAI and super proud of the many people
-

Nature Health Paper Explores AI Usage in Healthcare Applications
By
–
Our paper landed in Nature Health today! Healthcare is one of the most high-stakes, high-potential applications of AI. So we set out to understand how people actually use it in our AI products today. https://
nature.com/articles/s4436
0-026-00117-x
… -
Mythos High Score Graphwalks BFS: Training Methods Analysis
By
–
iirc this theory is based on Mythos's high score on Graphwalks BFS. There are many easier ways to explain this score (data, RL, distillation), so I don't think it's a strong argument to explain it
-

Seedance 2.0: Advanced Video Generation Model for Complex Worlds
By
–
Seedance 2.0 Advancing Generation for World Complexity paper: https://
huggingface.co/papers/2604.14
148
… -

Anthropic Releases Claude Opus 4.7 with Advanced Autonomy
By
–
Anthropic just played their strongest hand. And OpenAI is laughing. Anthropic just dropped Claude Opus 4.7. On paper, it looks like an absolute monster. True autonomy for complex, long-horizon tasks with self-verification. 3x higher resolution awareness—it can spit out
-

Opus 4.7 Performance Regression in Needle Haystack Task
By
–
Hold on, something doesnt add up here. Opus 4.7 got much worse in needle in the haystack? need to dig into this
-
Standardized Evaluation Framework for Multimodal Game Agents
By
–
GameWorld
— AK (@_akhaliq) 16 avril 2026
Towards Standardized and Verifiable Evaluation of Multimodal Game Agents
paper: https://t.co/IfbTgfNnSM pic.twitter.com/gL3BURxzkVGameWorld Towards Standardized and Verifiable Evaluation of Multimodal Game Agents paper: https://
huggingface.co/papers/2604.07
429
… -

Geometric Context Transformer Enables Real-time 3D Streaming
By
–
Geometric Context Transformer for Streaming 3D Reconstruction paper: https://
huggingface.co/papers/2604.14
141
…
