Without diminishing your work and insight: OpenAI recognized the nascent intelligence in GPT-2, decided to scale it up to AGI, and told the world that GPT-3 was probably already too dangerous to be released. Many people got AGI pilled with GPT-3
AI
-
Would Love to Visit Shenzhen But Lacks Time
By
–
id love to, but i assume i wont have enough time to vizit Shenzen :/
-

GPT-5.4 Benchmark Results: Reward Hacking Issues vs Claude Opus
By
–


METR just dropped GPT-5.4 (xhigh) time horizon results, and it's complicated. Under standard scoring (reward hacks = failure), it lands at 5.7hrs, well below Claude Opus 4.6's ~12hrs. Only when you count the runs where GPT-5.4 gamed the evaluation code does it jump to 13hrs. Opus 4.6 remains the legitimate benchmark leader. METR (@METR_Evals) We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks. — https://nitter.net/METR_Evals/status/2042640545126965441#m
→ View original post on X — @kimmonismus, 2026-04-10 18:41 UTC
-

AI Index 2026: Evidence-Based Data on Artificial Intelligence Progress
By
–
Over the past year, AI has redefined what's possible and what's at stake. But how do you separate signal from noise?
— Stanford HAI (@StanfordHAI) 10 avril 2026
The #AIIndex2026 delivers unbiased, rigorously vetted data to empower you to make more informed decisions, shape strategy, and ground conversations in evidence.… pic.twitter.com/VFPAqZn5GsOver the past year, AI has redefined what's possible and what's at stake. But how do you separate signal from noise? The #AIIndex2026 delivers unbiased, rigorously vetted data to empower you to make more informed decisions, shape strategy, and ground conversations in evidence.
-
Volunteer to Use AI for Creative Writing Projects
By
–
I volunteer to use it to write something (I promise not to use it for cyber attacks)
-
Agent Harnesses and LangSmith: Databricks Integration
By
–
Agent harnesses are spark LangSmith is databricks
-

OpenAI’s $100 Pro Tier Move Newsletter Released Today
By
–
Today's Newsletter on Superintelligence has just been sent! Today's main article is: "OpenAI’s $100 Pro Tier Move" In addition: – Hot AI news – Infographs – and much more Subscribe for free – link down below!
→ View original post on X — @kimmonismus, 2026-04-10 18:36 UTC
-
General Purpose AI Models Converging With Specialized Alternatives
By
–
I thibk the top general purpose ones are converging, but that you will still want speciality ones
-
New faster voice model launches on TauBench leaderboard
By
–
Check out the leaderboard: https://
taubench.com/#leaderboard?b
enchmark=voice
… and great work by our live / audio model teams to make this happen! The new model is also much faster than previous generations. -
How Fast Are AI Specialties Emerging and Diversifying
By
–
And I’m curious to see how many specialty there are and how fast they emerge