Mixture-of-Experts (MoE) models like DeepSeek-R1 unlock new levels of capability—but only if they can scale efficiently.
That’s where extreme hardware–software co-design at rack-scale comes in. With NVIDIA Blackwell and NVIDIA Dynamo, AI service providers can transform clusters
LLMS
-

MoE Models Scale Efficiently with NVIDIA Blackwell Hardware
By
–
-
Sparse Models Reveal Interpretable Task-Specific Components
By
–
Unlike with normal models, we often find that we can pull out simple, understandable parts of our sparse models that perform specific tasks, such as ending strings correctly in code or tracking variable types. We also show promising early signs that our method could potentially
-
New Interpretable Training Method for Small Language Models
By
–
We’ve developed a new way to train small AI models with internal mechanisms that are easier for humans to understand. Language models like the ones behind ChatGPT have complex, sometimes surprising structures, and we don’t yet fully understand how they work. This approach
-

GPT 5.1 launches on ChatGPT and surprises, but not as expected
By
–
GPT 5.1 just released on #ChatGPT and it shocked me… But not as you imagine → https://youtu.be/MPWedHYy9dw #GPT5 #GPT5_1
-
GPT5 Performance Issues and Reliability Concerns
By
–
J’espère qu’il va mieux fonctionner que GPT5 ! Parce que c’est la cata en ce moment.
-

Baidu Releases ERNIE 5.0 Preview
By
–

Baidu released ERNIE 5.0 Preview omnimodal foundational model. It is now available for testing on Ernie Chat.
-
ERNIE-4.5-VL-28B-A3B-Thinking Open-Sourced Under Apache 2.0
By
–
For developers and businesses, this is a major win: ERNIE-4.5-VL-28B-A3B-Thinking is fully open-sourced under Apache 2.0, meaning it’s ready for commercial use.
-
New Accessible High-Performing Multimodal AI Model Released
By
–
Integrate it, test it, or benchmark it — this model opens the door for more accessible, high-performing multimodal AI. Explore it here:
-

Baidu Open-Sources ERNIE-4.5-VL-28B Multimodal Model
By
–
Big news from @Baidu_Inc
! ERNIE-4.5-VL-28B-A3B-Thinking — a breakthrough lightweight multimodal reasoning model — is now officially open-sourced. It’s already trending #1 on Hugging Face’s Image-Text-to-Text models list. -
3B Parameter Model Matches GPT-5 and Gemini Performance
By
–
With just 3B active parameters, it matches Gemini-2.5-Pro and GPT-5-High, and even surpasses them on ChartQA and DocVQAval What can it achieve with only 3B active parameters?
For example, its “Thinking with Images” feature lets it zoom in, analyze details, and reason through