AI Dynamics

Global AI News Aggregator

About

Terminal-Bench 2 Scores Jump to 75-80% in 4 Months

Top scores on Terminal-Bench 2 went from ~25% → 75-80% in just 4 months. For Benchtalks #1, @vincentsunnchen sat down with @alexgshaw to dig into what happens when your benchmark gets solved before you're ready for the next one. Key takes: → The terminal is the right abstraction for agentic AI → Harbor exists because benchmarking and RL at scale are infra problems → "Benchmaxxing" is real; the defense is shipping harder tasks faster → TB3 is coming, and they want your hardest unsolvable problems "We need 1000x more benchmarks than we have right now" — @alexgshaw

→ View original post on X — @snorkelai, 2026-03-31 19:21 UTC