Spent most of the day on it, great coding and agent capabilities. I would choose this over any version of Gemini in any sort of agentic flow that requires deep coding understanding. Also, the way the model response looks is really interesting, it's very Claude-esque.
LLMS
-
WindSurf Google OpenAI Breakdown Video Analysis
By
–
Great breakdown video of what’s happening with WindSurf/Google/OpenAI!
-
Search-Based Reasoning: The Overlooked Path to AI Scaling
By
–
the success of: – o3 pro
– Grok 4 Heavy
– Gemini Deep Think is the single biggest sign to me that search is not dead as a path to scaling intelligence; we were just looking in the wrong place. -
AI Training on Books for Intellectual Discovery
By
–
In an academic bookstore and it is one of the times where I want a good AI trained on all books, even imperfectly. I want to learn a bit about the smells of antiquity & the history of idea of gray & etc. but am not going to read every book. I could learn a lot from an AI who has.
-
Achieving 100 Tokens Per Second Performance Benchmark
By
–
You should get at least 100 tokens/s on this bad boy!
-

Muon Algorithm Enables Scalable LLM Training Optimization
By
–
Muon is Scalable for LLM Training Liu et al.: https://
arxiv.org/abs/2502.16982 #ArtificialIntelligence #DeepLearning #MachineLearning -

New Dataset Release: Multilingual Reasoning Data Available
By
–
Incredible release! There are some gems in this dataset, like multilingual + reasoning data