5/5 Wrote up the complete investigation-from confident gibberish to the scheduler fix. Includes the false starts (suspected CUDA bugs), the breakthrough (request tracking), and lessons for anyone running stateful models at scale. Read the blog:
CODE
-
GPU Memory Utilization Bug: Token Classification Mismatch
By
–
2/5 Used gpu_memory_utilization=0.2 to reproduce quickly, but this happens naturally when the scheduler runs out of token budget and GPU blocks get recycled. New request gets 1 token → misclassified as "decode" But num_computed_tokens=0 → should be "prefill".
-

Debugging vLLM: Silent Corruption Bug in Jamba RL Training
By
–
1/5 Debugging vLLM: The silent corruption bug
1/1000 Jamba generations collapsed into confident gibberish during RL training. No crashes, no errors, just wrong outputs with high logprobs. The culprit? A scheduler edge case that only triggers under memory pressure. -

Custom Software Development at OpenAI with AI Tools
By
–
"Why do I buy planning software off the shelf?
Why wouldn't I code the kind of software that is exactly what OpenAl needs?" "OpenAl's finance team is just 18% the size of comparable companies. Why? Because their employees are building their own too with AI." "Building bespoke -
Prompt Ensembling Beats Single Prompts for High-Impact Decisions
By
–
When prompt ensembling beats single prompts: High-stakes business decisions ($100K+ impact) Medical/legal advice (where errors cost lives/money) Creative work (logo design, brand names, campaign ideas) Strategic planning (5-year roadmaps, market entry) Code
-
Style-conditioned AI models with prompt engineering capabilities
By
–
Oh yes. I liked the original Ai2 project for that reason, it's contained to a single codebase! So it can learn your style… I bet it'd be relatively easy to make it style-conditioned model as well, like a promptable feel/patterns/quality!
-

Implementing Genetic Algorithms for Optimization and Simulation
By
–
Hey coders and computational scientists! If you have never implemented Genetic Algorithms for your simulation, parameter search, and optimization problems & challenges, then you have really missed out. Get this book. I used GAs in my galaxy collisions simulation research as a
-

Claude Code vs ClawdBot: Security and Data Integration Comparison
By
–
I prefer using Claude Code directly over ClawdBot; getting the same value when I’ve it wired up to all my data. Getting security right is real and Claude Code has a lot of protections built in .. be safe out there if you’re trying out ClawdBot!
-
Google Adds CI Fixer and New MCP Integrations
By
–
Google released MCP integrations and a new "CI Fixer" feature in Beta for Jules SWE Agent.
— 🚨 AI News | TestingCatalog (@testingcatalog) 29 janvier 2026
CI Fixer "When enabled, Jules will automatically attempt to fix failing CI checks on pull requests it creates and push updates."
New MCPs include Linear, New Relic, Supabase, Neon,… pic.twitter.com/kKo8H9aGmCGoogle released MCP integrations and a new "CI Fixer" feature in Beta for Jules SWE Agent. CI Fixer "When enabled, Jules will automatically attempt to fix failing CI checks on pull requests it creates and push updates." New MCPs include Linear, New Relic, Supabase, Neon,
-
Using AI coding assistant to manage code changes
By
–
I just ask codex to remove the changes I don’t want and it works like magic