… and I didn't even mention DeepSeek! They've had a quiet couple of months, their last model release was the DeepSeek-R1-0528 family back in May
LLMS
-

Chinese AI Labs Release Major Models: Kimi, GLM, Qwen
By
–
July has been a truly incredible month for model releases from China – Moonshot (Kimi K2), http://
Z.ai (GLM-4.5) and 5 new releases from Qwen I think it's undeniable that the best available open weight models now come from the Chinese AI labs -
Neo: First Autonomous ML Engineer Agent Released
By
–
The first autonomous ML Engineer Agent has been released by @withneo! Neo is a multi-agent system that thinks, learns and builds like a real engineer.
— 🚨 AI News | TestingCatalog (@testingcatalog) 30 juillet 2025
– Perform analysis
– Finetune Llama and Gemma LLMs
– Prepare a visualisation
– Prepare training and evaluation pipelines https://t.co/1vvK4ddytw pic.twitter.com/PXN8WnRvYCThe first autonomous ML Engineer Agent has been released by @withneo
! Neo is a multi-agent system that thinks, learns and builds like a real engineer. – Perform analysis
– Finetune Llama and Gemma LLMs
– Prepare a visualisation
– Prepare training and evaluation pipelines -

Meta Developing Personal Superintelligence
By
–

BREAKING : Meta is building Personal Superintelligence! “Over the last few months we have begun to see glimpses of our AI systems improving themselves. The improvement is slow for now, but undeniable. Developing superintelligence is now in sight.”
-

MIT Research: How Language Models Track Dynamic Scenarios
By
–
How do language models track dynamic scenarios, like completing code or guessing your responses? MIT research finds that LMs don't track "state changes" step by step — they use mathematical shortcuts that we can control to boost LMs' prediction skills: https://
bit.ly/4lJfYfK -
Pelican Benchmark: Five-Attempt Model Selection Process
By
–
I've been meaning to do a version of the pelican benchmark where each model gets five attempts and then something picks the "best" one which moves on to the next round
-
Qwen3 Releases Five Major Model Updates in Nine Days
By
–
It's been a very busy 9 days for Qwen! Qwen3-235B-A22B-Instruct-2507 – 21st July
Qwen3-Coder-480B-A35B-Instruct – 22nd July
Qwen3-235B-A22B-Thinking-2507 – 25th July
Qwen3-30B-A3B-Instruct-2507 – 29th July
Qwen3-30B-A3B-Thinking-2507 – today -

Qwen3-30B Models Generate Pelican Bicycle Art Comparison
By
–
On the left the pelican on a bicycle from Qwen3-30B-A3B-Thinking-2507, on the right the pelican from Qwen3-30B-A3B-Instruct-2507
-

Qwen 3 30B Model Successfully Generates Functional Space Invaders Game
By
–
Not a great pelican, but it did write me a functional version of space invaders (unlike yesterday's non-thinking 30B-A3B which produced a game that didn't quite work) https://
simonwillison.net/2025/Jul/30/qw
en3-30b-a3b-thinking-2507/
… -
GPT-5 Code Capabilities Threaten Anthropic Revenue
By
–
What happens to that revenue if GPT-5 is much better at code? Cursor et al swith over & Anthropic revenue down 50% overnight?