Here's my detailed review – I put it through its paces, reverse engineered it a bit, then used it for some simple SQLite analysis followed by the same difficult US Census data chart recreation task that I posed to ChatGPT last night
@simonw
-
Claude Gets Code Interpreter: Sandboxed Python Node.js Execution
By
–
Anthropic are massively burying the lede here – they've called this "Upgraded file creation and analysis" (that really does seem to be the official name) but it's actually….
— Simon Willison (@simonw) 9 septembre 2025
Claude Code Interpreter!
It's sandboxed server-side Python/Node.js code execution for Claude https://t.co/yqsxWvRbY8Anthropic are massively burying the lede here – they've called this "Upgraded file creation and analysis" (that really does seem to be the official name) but it's actually…. Claude Code Interpreter! It's sandboxed server-side Python/Node.js code execution for Claude
-

ChatGPT Transcript Demonstrates Advanced Use Cases Beyond Manual Alternatives
By
–
I propose this ChatGPT transcript as an end-level boss for the "you could have done this just as easily without an LLM" crowd to take on https://
chatgpt.com/share/68bf48cf
-0e70-8006-a045-96fa8e7ddfc1
… -
Quantized Models Daytime: Specific Technical Accusations Explained
By
–
I've seen you say "1.58-bit quantized models during daytime" a few times now, that's a very specific accusation, where does it come from? Daytime in which timezone?
-
Building Compilers Costs 14k USD Across Three Languages
By
–
Re: costs Technically it costs about 5k usd to build your own compiler now because cursed was implemented first in c, then rust, now zig. So yeah, it’s not one compiler it’s three editions of it. For a total of $14k USD. x.com/GeoffreyHuntle…
-
Law’s Unpredictability: Judges and Strategic Legal Maneuvering
By
–
I'm beginning to appreciate that law is rarely as straight-forward as "it worked in this case so it will work in others too" – it all comes down to the individual judge (or sometimes jury), so most of the game is continually stacking the deck in your favor as much as possible
-
Fast mode default with quality fallback strategy
By
–
I tried it and it was noticeably slower so I'm defaulting to fast mode and switching up if it seems to have missed something important
-
GPT-5 Thinking Model Shows Significant Reliability Improvement
By
–
Are you seeing this with the new GPT-5 Thinking model that came out a few weeks ago? That's why I wrote about this now: I think it may have tipped over from completely untrustworthy to often (maybe even usually) correct
-
GPT-5 Shows Improved Accuracy Over Previous Versions
By
–
Reasonably carefully in this case – in the past I've seen all sorts of incorrect analysis and hallucinated details, GPT-5 seems much better on that front
-
Google’s New AI Mode: Impressive Alternative to AI Overviews
By
–
Follow-up note about a Google's new "AI mode" – it's actually very good! Massively different from "AI overviews" which is terrible