2/5 Used gpu_memory_utilization=0.2 to reproduce quickly, but this happens naturally when the scheduler runs out of token budget and GPU blocks get recycled. New request gets 1 token → misclassified as "decode" But num_computed_tokens=0 → should be "prefill".
COMPUTING
-
Rolv’s AI performance breakthrough proving itself with faster math
By
–
Months ago I shared Rolv with you and he shared that he had a major performance breakthrough for running AIs. His breakthrough is proving itself. Will share the final results when it's up, but his math is working out to speed things up.
-
Google Adds CI Fixer and New MCP Integrations
By
–
Google released MCP integrations and a new "CI Fixer" feature in Beta for Jules SWE Agent.
— 🚨 AI News | TestingCatalog (@testingcatalog) 29 janvier 2026
CI Fixer "When enabled, Jules will automatically attempt to fix failing CI checks on pull requests it creates and push updates."
New MCPs include Linear, New Relic, Supabase, Neon,… pic.twitter.com/kKo8H9aGmCGoogle released MCP integrations and a new "CI Fixer" feature in Beta for Jules SWE Agent. CI Fixer "When enabled, Jules will automatically attempt to fix failing CI checks on pull requests it creates and push updates." New MCPs include Linear, New Relic, Supabase, Neon,
-
National Quantum Initiative Reauthorization Strengthens US Computing Leadership
By
–
Quantum and AI are rapidly converging as a tool to advance science, security, and the global economy. Reauthorizing the National Quantum Initiative is how the U.S. can stay ahead and reinforce leadership in next-generation computing. Learn why the NQI matters now more than
-

Open-source Kimi 2.5 on own infrastructure with no limits
By
–
Damn, but open-source models like Kimi 2.5 on your own infrastructure WITH NO LIMITS AT ALL, but WHAT A BLAST !!!!! That's some crazy shit !!!!!
-

Edge AI Retail Solutions Showcased at EuroShop 2026
By
–
Heading to #EuroShop 2026 to share how edge AI enables more.
Thrilled to have our partner @InnowiseGroup joining us to demo deployable retail AI solutions!
If you're in Düsseldorf Feb 22-26, swing by or book a chat: https://
eu1.hubs.ly/H0rhM4C0
#RetailAI #EdgeAI #AIInference -
Cohere launches Model Vault for secure AI model deployment
By
–
Today we’re launching Model Vault — a dedicated, fully managed platform to run Cohere models securely and at scale.
— Cohere (@cohere) 28 janvier 2026
Model Vault delivers the control and isolation of self-hosting, without the operational burden:
🔒 Dedicated, isolated VPC
⚡ No noisy neighbors or rate limits
🔁… pic.twitter.com/ESkpISB99QToday we’re launching Model Vault — a dedicated, fully managed platform to run Cohere models securely and at scale. Model Vault delivers the control and isolation of self-hosting, without the operational burden: Dedicated, isolated VPC No noisy neighbors or rate limits
-

AXE Layout: Unified AI Hardware Optimization Across GPUs
By
–
What if one layout could rule all your AI hardware? Researchers from CMU, SJTU, NVIDIA, Princeton, and University of Toronto present AXE Layout. It's a single, simple map that tells data where to go across GPUs, memory, and threads—unifying sharding, tiling, and more.