First, let's talk about TPU 8t, which is designed for large-scale training and inference throughput. The pod size is increased slightly to 9600 chips, and provides ~3X the FP4 performance per pod vs. Ironwood (8t has 121 exaflops/pod vs. 42.5 exaflops/pod for Ironwood). In
RESEARCH
-

Young AI Researchers Win Award Without PhDs
By
–
Congrats to @AlecRad @Luke_Metz and @soumithchintala
! It's really cool to see work produced by a group of 20-somethings without a single PhD between them winning this award. I hope to see these three collaborate again one day 🙂 -

Unpredictable AGI may resist control, diverse AI safer
By
–
Unpredictable AGI may resist full control, making diverse #AI safer
by Gaby Clark @TechXplore_com Learn more: https://
bit.ly/48empCF #ArtificialIntelligence #MachineLearning #ML -

Anthropic Mythos vs GPT-5.5 CyberGym Benchmark Scores Compared
By
–
Also, a detail some of you may have missed: Anthropic’s gated Mythos model scored 83% on CyberGym. OpenAI just dropped GPT-5.5 at 82%. Except GPT-5.5 is actually available for people to use.
-
AI Philosophy Trend Growing More Painful Weekly
By
–
Yes, the new AI philosophy trend is growing more painful by the week
-

AI Models Automate Self-Jailbreaking Process Advancement
By
–
I mentioned this project in this week's AI Lab newsletter. It suggests a kind of flip side to the idea of self-improvement in AI models. As models become more powerful, it shows they can automate the process of figuring out how to jailbreak themselves and other models remarkably
-

GPT-5.5 performance on ARC-AGI-2 benchmark
By
–
GPT-5.5 just hit 85% on ARC-AGI-2. Somewhere, Yann LeCun is explaining why this still doesn't count. LLMs keep climbing a tall tree toward the moon
-
Claude 5.5 Outperforms Opus Beyond Frontend Work Analysis
By
–
Having used both extensively, this doesn’t track with my experience. Aside from frontend work, 5.5 is leaps and bounds ahead of Opus. I don’t think this benchmark tells us anything useful anymore.
-
Testing AI Model Performance Beyond Benchmarks Real Evaluation
By
–
Now it's time to test the model to draw clear conclusions not just from benchmarks, but from the real vibes of working with it. And also wait for external evaluations on the rest of the benchmarks from those who go on testing via the API. We'll keep you posted
-

ARC-AGI 5.5 outperforms Gemini 3.1 Pro at reasoning
By
–
In ARC-AGI 5.5, it is positioned at the cost frontier of Gemini 3.1 Pro but achieving a better score in its highest reasoning levels.
