🚨 STOP WHAT YOU ARE DOING AND LOOK AT THESE BENCHMARKS@atomic_chat_hq just unlocked 1.5x faster inference for Gemma 4 on LLaMA.cpp using Multi-Token Prediction.
— Charly Wargnier (@DataChaz) 8 mai 2026
138 tokens per second on a local 26B model is pure sorcery 👀
Get the code and GGUFs below ↓ https://t.co/o4aF64B5ee
STOP WHAT YOU ARE DOING AND LOOK AT THESE BENCHMARKS @atomic_chat_hq just unlocked 1.5x faster inference for Gemma 4 on LLaMA.cpp using Multi-Token Prediction. 138 tokens per second on a local 26B model is pure sorcery Get the code and GGUFs below ↓