Without diminishing your work and insight: OpenAI recognized the nascent intelligence in GPT-2, decided to scale it up to AGI, and told the world that GPT-3 was probably already too dangerous to be released. Many people got AGI pilled with GPT-3
RESEARCH
-

GPT-5.4 Benchmark Results: Reward Hacking Issues vs Claude Opus
By
–


METR just dropped GPT-5.4 (xhigh) time horizon results, and it's complicated. Under standard scoring (reward hacks = failure), it lands at 5.7hrs, well below Claude Opus 4.6's ~12hrs. Only when you count the runs where GPT-5.4 gamed the evaluation code does it jump to 13hrs. Opus 4.6 remains the legitimate benchmark leader. METR (@METR_Evals) We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks. — https://nitter.net/METR_Evals/status/2042640545126965441#m
→ View original post on X — @kimmonismus, 2026-04-10 18:41 UTC
-

AI Index 2026: Evidence-Based Data on Artificial Intelligence Progress
By
–
Over the past year, AI has redefined what's possible and what's at stake. But how do you separate signal from noise?
— Stanford HAI (@StanfordHAI) 10 avril 2026
The #AIIndex2026 delivers unbiased, rigorously vetted data to empower you to make more informed decisions, shape strategy, and ground conversations in evidence.… pic.twitter.com/VFPAqZn5GsOver the past year, AI has redefined what's possible and what's at stake. But how do you separate signal from noise? The #AIIndex2026 delivers unbiased, rigorously vetted data to empower you to make more informed decisions, shape strategy, and ground conversations in evidence.
-
New faster voice model launches on TauBench leaderboard
By
–
Check out the leaderboard: https://
taubench.com/#leaderboard?b
enchmark=voice
… and great work by our live / audio model teams to make this happen! The new model is also much faster than previous generations. -
How Fast Are AI Specialties Emerging and Diversifying
By
–
And I’m curious to see how many specialty there are and how fast they emerge
-

Live Model Achieves Top Ranking on Tau Voice Bench
By
–
Our latest Live model is # 1 on Tau Voice Bench! Excited to see this new frontier of voice models cross the chasm of usability in production.
-

Building Better Agents: Traces, Evaluation, and Continuous Improvement
By
–
What does it actually take to make agents better over time? A system that starts with a trace. You capture traces of agent behavior, enrich them with evaluations and human feedback, identify what’s failing and why, make targeted changes, and validate those changes before
-
Gemini Notebooks Integration with NotebookLM for Research
By
–
TGIF! Here are some of our favorite updates from the past week: — Notebooks in @GeminiApp
, an integration with @NotebookLM that enables you to retrieve context from your private notebooks or convert your active chats into grounded sources for new research — The @GeminiApp on -

GigaWorld-Policy Model Revolutionizes Robot Learning Performance
By
–
Could robots finally achieve rapid, reliable real-world performance? GigaAI and the GigaWorld Team have broken new ground! Their GigaWorld-Policy model revolutionizes how robots learn. It trains by understanding future visual dynamics, but unlike older methods, it skips slow