Benchmarks are great, but IMO the behavior change is a much bigger deal. Plans before it edits, recovers from its own errors, and finds creative ways around obstacles instead of stalling. Feels much more like a senior engineer than 4.7, and better at long-horizon work.
RESEARCH
-

Claude Opus 4.8 achieves 69.2% on SWE Bench Pro, beating Opus 4.7
By
–

ANTHROPIC : Claude Opus 4.8 achieves 69.2% score on SWE Bench Pro against 64.3% for Opus 4.7. Benchmarks
-
AI Agents Need to Learn from Executions, Not Just Complexity
By
–
Exacto. Ese es el problema que nadie estaba atacando. Todos haciendo agentes más complejos pero ninguno que realmente aprenda de sus propias ejecuciones
-
AI Self-Improving Agent Outperforms Karpathy’s Autoresearcher
By
–
Esa gráfica es clave. Se ve claramente el momento en el que el self-improving agent se despega del autoresearcher de Karpathy y ya no baja
-

Linux Foundation OpenMDW Framework for Open Models
By
–
We're adopting the Linux Foundation’s OpenMDW framework across our open model families. This helps make open model licensing simpler and more consistent at scale. A single legal framework across models, code, documentation, and data helps reduce friction for developers and
-
AI Agents Need to Learn from Executions for Meaningful Functionality
By
–
I've been waiting for something like this for a while. Agents that don't learn from their executions make no sense.
-
Self Improving AI beats Karpathy’s autoresearcher by improving itself
By
–
Self Improving AI (SIA) beats Karpathy's autoresearcher agent by improving itself!
— Sumanth (@Sumanth_077) 28 mai 2026
SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model / Agent) on a benchmark task.
Most agent frameworks are static. Fixed harness, fixed model… https://t.co/bk4BaLqsCE pic.twitter.com/kiByM7J0KvSelf Improving AI (SIA) beats Karpathy's autoresearcher agent by improving itself! SIA is a Self Improving AI framework to autonomously improve the performance of any AI system (Model / Agent) on a benchmark task. Most agent frameworks are static. Fixed harness, fixed model
-

AutoScientists: Decentralized AI Agents for Scientific Research
By
–
Banger paper from Harvard. AutoScientists drops the central planner entirely. Agents interpret shared experimental data, self-organize around promising directions, evaluate proposals before resource allocation, and document successes AND failures. Decentralized AI co-scientists
-
MIT Deep Learning Course: Breakthroughs in AI Applications
By
–
Free MIT course breaks down deep learning: https://t.co/S9DmjSlveB
— MIT CSAIL (@MIT_CSAIL) 28 mai 2026
Here, MIT ass't prof. Sara Beery discusses how it's driven breakthroughs in areas like image generation, coding, & playing games (Lecture 1). pic.twitter.com/qXuJiI3MamFree MIT course breaks down deep learning: https://
bit.ly/4t9gHJL Here, MIT ass't prof. Sara Beery discusses how it's driven breakthroughs in areas like image generation, coding, & playing games (Lecture 1). -
Generative Supervision for Embodied Intelligence
By
–
GEM
— AK (@_akhaliq) 28 mai 2026
Generative Supervision Helps Embodied Intelligence pic.twitter.com/IlGPbxkwHSGEM Generative Supervision Helps Embodied Intelligence