probably 10x more people should be working on prompt optimization systems (we need a vLLM for promptopt), theory, new techniques, benchmarks. the whole kit and caboodle
LLMS
-
Context Window Performance and RAM Requirements for AI Models
By
–
60 tokens a second, context window is supposedly up to 262,144 but that will use a ton more RAM, probably more than my Mac can handle (I've not tried though)
-
Qwen3-Coder-30B Demonstrates Strong Capabilities Without Thinking
By
–
For a non-thinking model I'm finding Qwen3-Coder-30B to be surprisingly capable Thinking models are a pretty new invention, up until about 8 months ago we didn't have any at all
-
Hosted AI Models Approaching Quality Parity with Alternatives
By
–
They're getting pretty close, if they're not there already – depends on how sensitive you are to differences in quality, the hosted models are still more capable
-
LLM-assisted coding becoming standard developer practice within years
By
–
I'm confident that within a year or two it won't make sense to have a buzzword for "used an LLM to help me write code as a software dev" because that would be like having a buzzword for "searched Google and copied and pasted snippets from Stack Overflow" – that's just coding!
-
Using LLM-Generated Code Without Review: Naming the Practice
By
–
That's what I want it to mean, yes – I think it's really useful to have a name for "I got an LLM to write code that I just used without even reviewing it or caring about the code"
-
Running Qwen3-Coder-Flash 30B on M2 MacBook Pro
By
–
Here's the pelican riding a bicycle and the implementation of space invaders I got from running the 24.82GB 6bit MLX version of Qwen3-Coder-30B-A3B-Instruct – aka Qwen3-Coder-Flash – on my 64GB M2 MacBook Pro
-
Running Qwen3 Coder on Mac with LM Studio and Open WebUI
By
–
I wrote about the new Qwen3 Coder model here, including details on how I got it running on my Mac using LM Studio, Open WebUI, mlx-lm and my own LLM command-line tool https://
simonwillison.net/2025/Jul/31/qw
en3-coder-flash/
… -
Natural Birth Outperforms Multibillion-Dollar LLM Models
By
–
You can make a baby for free, and the result is still better than a multibillion-dollar LLM.
-

Voice as the Original and Future Modality of Communication
By
–
"Voice is our past as much as it is our future" Although roon's "text is the universal interface" proved true for LLMs, I do think voice is the "OG modality" for communication between intelligent species: – Before writing was invented, we were speaking for over 100,000 years.
–