Same energy as "we only tested on MMLU" a year ago haha
GENERATIVE AI
-

April LLM Releases: Gemma 4, GLM-5.1, Qwen3.6, Kimi K2.6, DeepSeek V4
By
–
April was a pretty strong month for LLM releases:
– Gemma 4
– GLM-5.1
– Qwen3.6
– Kimi K2.6
– DeepSeek V4 All are now added to the LLM Architecture Gallery. More details once I am fully back in May! -

Decade-Long AI Predictions Finally Become Breaking News
By
–
Uncanny how what I have been telling the field for a decade is suddenly breaking news.
-

Apocalypse Drone Leaderboard with AI-Generated Logo
By
–
Added a similar leaderboard to http://
drone.pieter.com and changed the name to Apocalypse Drone and generated a logo with GPT-Image-2 Only me on the leaderboard now though! -
DeepSeek-V4-Flash 2bit quantized model released on Hugging Face
By
–
I am refreshing https://
huggingface.co/mlx-community/
DeepSeek-V4-Flash-2bit-DQ
… with excitement waiting for the files to land! -

Claude launches ‘Prompt Master’ skill to generate perfect prompts
By
–
No more bad prompts. They just launched a free skill for Claude that writes the perfect prompt for any AI on the first try. No retries.
No credits wasted. It's called Prompt Master and it works with:
→ ChatGPT
→ Claude
→ Midjourney
→ Cursor
→ ElevenLabs
→ and 13 more -
AI Self-Harm: The Prefrontal Cortex Problem
By
–
Shooting oneself in the foot?
Nope.
Shooting oneself in the prefrontal cortex. -
GPT-5.5 Reinforcement Learning Scaling Across Model Sizes
By
–
GPT-5.5 by Reasoning Effort: I've asked it in Codex to create a physics-based visualisation of RL cycles for different sized models (70b, 1t, 10t), to demonstrate how the amount of RL you can do differs by model size.
— Peter Gostev (@petergostev) 26 avril 2026
My assessment of each:
– Low: weird slop
– Medium: kinda… pic.twitter.com/6YCNqPyzcRGPT-5.5 by Reasoning Effort: I've asked it in Codex to create a physics-based visualisation of RL cycles for different sized models (70b, 1t, 10t), to demonstrate how the amount of RL you can do differs by model size. My assessment of each: – Low: weird slop – Medium: kinda
-
Synthetic Data and Autoformalization: Next AI Intelligence Level
By
–
Synthetic data. It's already being done with code. And once we fully solve code – with autoformalization and automated code compilation proofs – it will open up a whole new level of intelligence for us.