The latest crop of models remains below 1% on ARC-AGI-3 — for now. Where will the scores be by the end of the year?
RESEARCH
-

MIT Lab Technology Turned Into At-Home Product Using AI
By
–
#AI helped me turn MIT lab #Technology into an at-home product for parents
by Kara Baskin @MITSloan Learn more: https://
bit.ly/4t4McoD #ArtificialIntelligence #ML #MachineLearning #Tech -

Speculative Decoding Accelerates RL Rollouts 2.5x in NeMo-RL
By
–
RL post-training is hitting a rollout bottleneck. This new paper from #NVIDIAResearch shows how speculative decoding in NeMo-RL + @vllm_project can accelerate rollouts losslessly, with 1.8x higher throughput at 8B and projected 2.5x end-to-end speedup at 235B. Read the full
-

RL Boosts Known Tasks But Causes Hallucinations on Unknown Ones
By
–
RL is a bit of a double edged sword: in known territory performance increases, but in unknown territory the model tends to hallucinate that it is performing a completely different task it was trained on
-
Why Measure AI Performance With Humans Rather Than Alone
By
–
This is a laudable goal, but it may mean designing, testing, and deploying AI systems differently. For example, why not measure the ability of GPT-5.5 to solve problems *with* a person rather than on its own? (see Stanford's centaur
-

Survey Proposes Taxonomy for Feed-Forward 3D Reconstruction
By
–
What if you could reconstruct 3D from 2D in a single forward pass, no per-scene optimization needed? Researchers from Zhejiang University, NTU Singapore, Monash, ETH Zurich & Uni Tübingen present a new survey on feed-forward 3D reconstruction. They propose a taxonomy that
-

Why Compiling Code Doesn’t Mean Correct AI-Generated Software
By
–
“Marcus’ specific point about coding is structurally important: a model that produces code which compiles and passes the tests it was given is not the same as a model that produces correct, secure, maintainable, well-architected software. The first is verifiable in seconds; the
-

Microsoft Research Paper on Training Computer-Use Agents
By
–
NEW paper from Microsoft Research. If you care about training computer-use agents, this is one to keep. (bookmark it) The team builds 1,000 synthetic computers (each with realistic directory structures, documents, and artifacts) then runs long-horizon simulations on top of
-

Co-Evolving Policy Distillation New AI Research Paper
By
–
Co-Evolving Policy Distillation paper: https://
huggingface.co/papers/2604.27
083
…

