MRCR is a bad eval, we use GraphWalks now. More info here:
RESEARCH
-

Qwen 3.6-35B-A3B vs Claude Opus 4.7 SVG Generation Comparison
By
–
Here's Qwen 3.6-35B-A3B v.s. Claude Opus 4.7 for "Generate an SVG of a flamingo riding a unicycle", in case you thought Qwen might be cheating at the pelican benchmark
-
AI Transforming Healthcare Through Comprehensive Patient Data
By
–
Beautifully written piece by @FAbnousi about how AI for health might look like in the future
The current data in health is limited because it only captures episodic clinical snapshots of what happens to our bodies
The revelation is that there is so much latent knowledge in -
Letting Models Think: From Manual Temperature Control to Automatic Decision-Making
By
–
Better to let the model think. It's sort of like how we used to manually set "temperature" two years ago — nowadays it's better to let the model decide.
-
Mid and Post Training Compute Now Rivals Pretraining Cost
By
–
i think in a world where mid and post training take equivalent or more compute (!!) than pretraining, this is less big a deal than what it used to mean in 2023
-

Local Qwen 35B Outperforms Claude Opus 4.7 on Benchmark
By
–
Shocking result on my pelican benchmark this morning, I got a better pelican from a 21GB local Qwen3.6-35B-A3B running on my laptop than I did from the new Opus 4.7! Qwen on the left, Opus on the right
-

Phasing Out Flawed Evaluation with System Card Caveat
By
–
This is a bad eval that we've been phasing out. Going to add a caveat to the system card to make it clear. More here:
-

MIT CSAIL Harvard AI Accelerates Electron Microscope Imaging Tenfold
By
–
A decade of imaging.
— MIT CSAIL (@MIT_CSAIL) 16 avril 2026
Compressed into three months.
Here’s how MIT CSAIL & Harvard taught an electron microscope to see like you do: https://t.co/qC3zVck0HF pic.twitter.com/kMTptacRMDA decade of imaging. Compressed into three months. Here’s how MIT CSAIL & Harvard taught an electron microscope to see like you do: https://
bit.ly/4tfPxBX -

MRCR Phase-Out: Shifting from Distractor-Based to Applied Long-Context
By
–
We kept MRCR in the system card for scientific honesty, but we've actually been phasing it out slowly. Two reasons: (1) it's built around stacking distractors to trick the model, which isn't how people actually use long context, and (2) we care more about applied long-context
-

Phasing Out MRCR: Shifting Focus from Distraction Tricks to Applied Long Context
By
–
We kept MRCR in the system card for scientific honesty, but we've actually been phasing it out slowly. Two reasons: (1) it's built around stacking distractors to trick the model, which isn't how people actually use long context, and (2) we care more about applied
