I particularly appreciate how this rationale isn't based on the idea that LLM code is of poor quality compared to code written by hand – the quality of the code isn't the deciding factor at all here
LLMS
-
Low-Background Tokens and Vintage Language Models for Quality AI
By
–
> be me
— swyx 🇸🇬 (@swyx) 30 avril 2026
> "the internet is polluted by ai slop, we need low-background tokens"
> "wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us"
> all the existing vintage models are like <4B
> we need a chat tuned 13B vintage model
>… https://t.co/DcXA62Jjnd> be me
> "the internet is polluted by ai slop, we need low-background tokens"
> "wouldnt it be cool if we could time travel and see what our ancestors 100 years ago would say to us"
> all the existing vintage models are like we need a chat tuned 13B vintage model > -
Base Model Release Sparks Experimentation With API Completions
By
–
i havent done the work to compare it to peers but i'm just excited that we have a base model and honestly for all the people that complained about the death of the completions API (
@deepfates ? or deepfates adjacent) not enough people are experimenting with weird usages and -

Maestro Approach Achieves SOTA on Browsecomp-Plus
By
–
Live from @DeepLearningAI conference: our CPO Or Dagan is taking the stage, explaining how we got SOTA on Browsecomp-Plus with the Maestro approach.
-
Copilot Limitations Compared to Competing AI Solutions
By
–
Ok, to be fair, they can probably only use copilot in their company, and it's pretty bad compared to the other options, but still, wtf?
-
BioMysteryBench: Claude’s Creative Bioinformatics Research Solutions
By
–
BioMysteryBench, our new bioinformatics eval, tests whether Claude can devise creative solutions to open-ended research problems. Read more:
-

Claude Solves 30% of Expert-Stumping Biological Data Problems
By
–
New on the Science Blog: We gave Claude 99 problems analyzing real biological data and compared its performance against an expert panel. On 23 problems, the experts were stumped. Our most recent models solved roughly 30% of those—and most of the rest.
-

Gemini 3.1 Pro vs GPT-5.5 Pro: Ethics Impact on AI
By
–
Its a bit frustrating, because Gemini 3.1 Pro is an excellent model and can deliver really good results. But here is GPT-5.5 Pro for comparison. Sadly, it took this assignment very seriously and ethically and thus was no fun at all.
-
Claude’s measurable strengths: iteration, exploration, filtering
By
–
Claude’s superpower is measurable in output:
speed of iteration, number of paths explored, and how fast bad ideas get filtered out. Same model, different usage completely different power.