I am still playing with Agent responses, which are powered by the same Gemini 2.5 Pro, but the output feels more reliable than what we have on personal Gemini. Full scoop
LLMS
-
Qwen 3 release questioned technical report limitations
By
–
I think the only one I might question the significance of in your list is Qwen 3. They didn't even release a base model. The technical report was quite disappointing as well.
-
vLLM Limitations in Broader AI Landscape
By
–
vLLM is a small part of the whole picture. They're massively lacking.
-
Comparing Sonnet and Opus AI Models: Performance Differences
By
–
Are you using Sonnet or Opus? And when you say dumber what specifically?
-
Rate Limits and Token Compaction Bugs in API Systems
By
–
If you're seeing lower rate limits than what we publish, or if you're seeing auto-compact earlier than ~155k tokens, that's a bug. /bug to report.
-
Are We Near AGI Flight? Current LLM Capabilities Assessed
By
–
My short essay on LLMs & AGI: Have we reached the 'AGI of flight'? Not even close. Our planes can carry hundreds of passengers across the globe in just a few hours, yet nothing we build matches a godwit flying non-stop for eleven days without rest or the agility of a bat
-
Claude and GPT-5 for coding: shifting away from Cursor agents
By
–
I primarily use Claude with Claude code and GPT-5 in Codex. I rarely use cursor agents anymore.
-

Samsung’s Tiny AI Model Outperforms DeepSeek and Gemini
By
–
5. Samsung’s Tiny AI beats the giants Samsung unveiled the Tiny Recursion Model (TRM), a 7M-parameter AI that outperformed DeepSeek R1 and Gemini 2.5 Pro on reasoning tasks using a self-improvement loop.
-
LLaMA 2 Sparked Open Source LLM Interest
By
–
It definitely kickstarted things for LLMs but it was LLaMA 2 that showed there was interest in opensource models.
-
Qwen 2.5 influence research adoption compared to other models
By
–
not that it's a bad model, but in terms of influence that is yet to be determined Qwen 2.5 is still used to this day in research and experiments, more than any other model – that's a different level of influence