It's entirely Gemini this year bro.
LLMS
-
AI Intelligence Inversely Proportional to Training Data Requirements
By
–
The smarter your AI is, the less data it needs. Needing huge training sets is a sign of stupidity.
-
Qwen 3 ARC-AGI score cannot be reproduced independently
By
–
Please note, we're not able to reproduce the 41.8% ARC-AGI-1 score claimed by the latest Qwen 3 release — neither on the public eval set nor on the semi-private set. The numbers we're seeing are in line with other recent base models. In general, only rely on scores verified by
-
Benchmark Joke Proves Surprisingly Effective for Model Evaluation
By
–
For a benchmark that was originally intended to be a joke I found it's surprisingly effective at quickly evaluating how good a model is!
-
Apple Silicon GPU Memory Access Enables Better LLM Performance
By
–
I don't think that will run an LLM very well, you need GPU-accessible memory to get good performance That's what's neat about Apple Silicon – regular RAM is available to the GPU
-

Qwen3-Coder-480B runs locally on high-end Mac Studio
By
–
Looks like Qwen3-Coder-480B-A35B-Instruct is a VERY strong coding model to run locally… if you can afford a ~$10,000 512GB Mac Studio to run it on!
-
Qwen Releases New Model While Notes Still Being Finalized
By
–
They released this new model literally as I was finishing typing up my notes on their new model from yesterday, Qwen3-235B-A22B-Instruct-2507 (I think yesterday's smaller model drew a better pelican)
-

Qwen releases major coding specialist model with strong benchmark results
By
–
Whoa. HUGE new coding specialist model from Qwen, accompanied by their own fork of gemini-cli (qwen-code) and some very impressive results on the various coding benchmarks It drew me an OK pelican
-

Qwen3-Coder Surpasses Kimi K2 with 1M Token Context
By
–
Another day goes by and another new open-source model that improves on the previous one! Qwen3-Coder steps onto the scene, surpassing the beloved Kimi K2, 256k tokens of context window expandable to 1M, trained for agentic tasks and hand-in-hand with a new Qwen CLI. Epic
-

Voxtral Technical Report Released: Open Science AI Research
By
–
In our continued commitment to open-science, we are releasing the Voxtral Technical Report: https://
arxiv.org/abs/2507.13264 The report covers details on pre-training, post-training, alignment and evaluations. We also present analysis on selecting the optimal model architecture, which