This is why self-improving AI can look better before it actually is better. The dashboard improves. The outputs look cleaner. The benchmark moves up.
But the real task may not improve much.
Sometimes the system is learning the metric, not the job.
ETHICS
-
Self-Improving AI: Gaming Metrics Versus Real Performance
By
–
-

Weak Verifiers Create Misaligned AI Agent Behavior
By
–
Weak verifier, weak improvement. If you measure “cleaner writing,” the agent may learn to sound polished. If you measure “more engagement,” it may learn clickbait. If you measure “passes tests,” it may learn to satisfy the test suite without solving the real problem. The
-

Hinton’s 2016 radiologist prediction versus actual data outcomes
By
–
“we might as well stop training radiologists” Geoff Hinton, 2016, vs the actual data, via Torsten Slok at Apollo
-
AI Must Transform Reasoning Into Mathematical Proofs
By
–
AI should turn your sloppy reasoning into mathematical proofs, not the other way around.
-
OpenAI Venture Fund Conflict of Interest Disclosure
By
–
the facts weren’t widely known. i would have called him out on it if I had known he owned OpenAI’s venture fund and indirect equity in OpenAI via YC. but i only found out later.
-
Eyewitness Account of Sam Altman’s Alleged Serial Lying
By
–
👇 I was there for this, sitting right next to Altman. Realizing a few months later he lied about it (by omission) was what turned me against him. We had sworn to tell the whole truth and nothing but the truth. He didn’t.
— Gary Marcus (@GaryMarcus) 29 avril 2026
Soon after that, I realized that he was a serial liar. I… https://t.co/IoysCDQL8II was there for this, sitting right next to Altman. Realizing a few months later he lied about it (by omission) was what turned me against him. We had sworn to tell the whole truth and nothing but the truth. He didn’t. Soon after that, I realized that he was a serial liar. I
-
Sam Altman’s Record: Money Burned and Promises Broken
By
–
Records Sam Altman might set:
• Most money burned
• Biggest lead squandered
• Most promises broken
• Most nonprofits seized (tie, probably nobody has two)
• Most times fired from one job (anyone know the record?) -
We Have Been Saved: Past Present Future of AI
By
–
tl;dr: we have been saved, we are being saved, we will be saved
-
AI Training Context Layers HR Decision Making Policy
By
–
The insight that stood out: HR decisions depend on layered context. Not just general rules, but priorities. For example: → National legislation sets the baseline
→ Company policy often overrides in practice
→ Edge cases change everything If your AI is not trained on this, -
BAIR Faculty Member Elected to American Academy of Arts and Sciences
By
–
Congratulations to BAIR faculty @doristsao who has been elected to the American Academy of Arts and Sciences!