Enjoyed the read? If you have deep experience in ML frameworks (training or inference) and love working on problems like these, our team is hiring! ML Systems Engineer, Frameworks & Tooling: https://
jobs.ashbyhq.com/cohere/c99e61c
9-ed92-426d-9711-188dfc0f729f?departmentId=7130c75e-15b8-493c-959f-e9b8ea5c1c09
… Audio Inference Engineer, Model Efficiency:
@cohere
-
Cohere Hiring ML Systems and Audio Inference Engineers
By
–
-

BF16 to FP8 Quantization: Per-Channel Scaling for LLM Accuracy
By
–
The tricky part: naïvely casting BF16 group scales to FP8 dropped the quality. Our fix: quantize scales per-channel (outer vector scaling) + rescale by 1/8 to avoid FP8 clipping. Result: >99.5% of W4A16 accuracy recovered on Command A & Cohere MoE. Paired with a CUTLASS
-

W4A8 Inference Production-Ready Integration in vLLM
By
–
Excited to share our work on production-ready W4A8 inference, now integrated in vLLM! By combining 4-bit weights (low memory) with 8-bit activations (high compute), we hit the sweet spot for both decoding and prefill — up to 58% faster TTFT and 45% faster TPOT vs W4A16 on Hopper.
-
Speculative Decoding Optimization for Mixture of Experts Models
By
–
Get more from speculative decoding in MoE models
-
MoE-Based LLMs Enhance Speculative Decoding Effectiveness
By
–
New Technical Report from @EkagraRanjan
: Contrary to what you might expect, MoE-based LLMs make speculative decoding even more effective. Read more on our blog: -

Company Named Forbes AI 50 for Secure Sovereign AI
By
–
We're proud to be on the @Forbes AI 50 list again! It reflects our focus on building secure, sovereign AI that helps enterprises and governments put their data to work on their terms. Learn more: https://
forbes.com/lists/ai50/ -
RWS Language Weaver AI Translation Model Powers Command
By
–
We worked with @RWSGroup to fine-tune our Command translation model, which improved language and cultural expertise. Now, that model is the “brain” that powers RWS’ Language Weaver AI translation solution. https://t.co/0xCCwG3EU0
— Cohere (@cohere) 10 avril 2026We worked with @RWSGroup to fine-tune our Command translation model, which improved language and cultural expertise. Now, that model is the “brain” that powers RWS’ Language Weaver AI translation solution.
-

Cohere and Microsoft discuss enterprise AI acceleration on Model Mondays
By
–
How can enterprises accelerate decision-making and improve efficiency? We're joining @Microsoft's Model Mondays on April 6th to discuss how AI can help you: 🌎 Drive real-world enterprise value 💡 Surface precise, verifiable insights 📌 Use Microsoft Azure and Cohere enterprise models to accelerate your AI journey
-
Cohere Unveils Custom LLM for Healthcare Management
By
–
Learn more: cohere.com/blog/ensemble-cohere-custom-healthcare-llm-for-rcm [Translated from EN to English]