It actually holds up on real reasoning tasks like GSM8K and Countdown (not toy stuff) and matches GRPO pretty well. They even did full pretraining of a recurrent LM from scratch in pure int8. Sample efficiency is lower than backprop (normal for evolution strategies), but you get
LLMS
-
Google’s AI Agents Build Working Operating System in 12 Hours
By
–
Google's Antigravity 2.0 built the core framework of a working operating system in 12 hours, spinning up 93 sub-agents and processing billions of tokens for under $1,000 in compute costs.
— Chubby♨️ (@kimmonismus) 20 mai 2026
On stage, the team booted Doom on the AI-built OS.
Just imagine all the possibilities in… pic.twitter.com/yT2RffsjrwGoogle's Antigravity 2.0 built the core framework of a working operating system in 12 hours, spinning up 93 sub-agents and processing billions of tokens for under $1,000 in compute costs. On stage, the team booted Doom on the AI-built OS. Just imagine all the possibilities in
-

Developers build RAG and multi-agent AI pipelines on Google Cloud
By
–
At #GoogleIO, NVIDIA and @GoogleCloud are celebrating a milestone: More than 100,000 developers have joined our joint developer community in just one year, and the platform is expanding. In the past year, members have shipped RAG applications on GKE, built multi-agent pipelines,
-
Release of Nemotron-Labs-Diffusion parallel generation language models
By
–
Most language models only generate one token at a time.
— NVIDIA AI (@NVIDIAAI) 19 mai 2026
We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather than committing to each token permanently,… pic.twitter.com/fTOBmQ8KaMMost language models only generate one token at a time. We just released Nemotron-Labs-Diffusion, a family of diffusion language models that take a different approach, generating multiple tokens in parallel within a single model. Rather than committing to each token permanently,
-
Major AI platforms converging or diverging: who will win?
By
–
The gap between what you can do on ChatGPT/Codex and Claude/Code/Cowork is closing, as Anthropic & OpenAI converge on a single experience. Google's experiences are diverging: Studio & Gemini & Antigravity & the other Google AI apps are increasingly different. Which will win?
-

Gemini 3.5 Flash performance in BullshitBench
By
–
BullshitBench update: Gemini 3.5 Flash did pretty badly – below a bit Gemma 4 even (31b high is 80th)
-

Antigravity provides better end-of-task transparency than Codex
By
–
Fascinatingly Antigravity is actually the best tool so far at providing this sort of transparency, doing something by default that Codex and Code do not: offering a summary of exactly what it did at the end of a task. Just add this to Gemini! (But also cite sources more)
-

Antigravity reveals model thinking traces
By
–
The crazy thing is that Antigravity does this quite well! So it isn't like Google is hiding these thinking traces because of distillation or because they are full of insane mutterings. I guess you need to either do the work in Antigravity or not at all?
-

The Evolution of Open-Source Large Models and Accessibility
By
–
I remember when people were saying "It's useless to open-source big models because nobody will be able to run them fast"….
-

User critique: Gemini models not ready for enterprise
By
–

I find this continually frustrating. Gemini 3.5 Flash is excellent, as is Gemini 3.1 Pro. But you absolutely cannot use them for any serious purpose right now, especially for any enterprise work. Compare to Claude or ChatGPT: you can understand what the model did & how to correct