Weird, what issue are you having. Works for me both in ollama and native llama.cpp I am using unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:MXFP4_MOE
LLMS
-

Honeywell RAG Agent for Field Technician Equipment Diagnostics
By
–
.
@Honeywell built a RAG-powered agent that helps field technicians diagnose and fix equipment on-site, running fully air-gapped on-prem. Alpesh Desai will break down how they built it at Interrupt. https://
interrupt.langchain.com -

GENIUS Suite Evaluates Multimodal AI Reasoning Capabilities
By
–
Can your favorite AI truly reason on the fly, or just remember what it's learned? A team from Peking University, CUHK, StepFun, PolyU, and MSRA introduces GENIUS, a groundbreaking evaluation suite. GENIUS challenges Unified Multimodal Models (UMMs) on Generative Fluid
-
Building Large Language Models From Scratch YouTube Series
By
–
I actually have a Build an LLM From Scratch YouTube series
-

2026 Keynote Pregame: Accelerated Computing, Open Models, Agentic AI
By
–
Your 2026 keynote pregame hosts are here. Join @saranormous (Conviction), @GavinSBaker (Atreides), @Alfred_Lin (Sequoia Capital), and Tiffany Janzen (TiffinTech) as they set the stage for fast-paced conversations on accelerated computing, open models, agentic AI, and the
-
US Decision-Makers’ Understanding of Generative AI and Intelligence Questioned
By
–
The people making decisions about AI in the US really don’t seem to understand how the generative AI models work or what intelligence is, or how to evaluate it.
-
Debunking LLM Consciousness and Supply Risk Claims
By
–
Wild.
— Gary Marcus (@GaryMarcus) 12 mars 2026
If any of this is true (and most if it isn’t*) it’s true of all LLMs and not just Anthropic’s models.
There is no technical difference between models that would make one more of a supply risk than another.
*mimicking text doesn’t make an AI conscious or anxious etc. but… https://t.co/SzI5FULryYWild. If any of this is true (and most if it isn’t*) it’s true of all LLMs and not just Anthropic’s models. There is no technical difference between models that would make one more of a supply risk than another. *mimicking text doesn’t make an AI conscious or anxious etc. but
-

Scalable MoE Training Efficiency with Megatron Core
By
–
“Scalable Training of Mixture-of-Experts Models with Megatron Core” This NVIDIA MoE report walks through the hard part of MoE training. The key is not to add more parameters, but keeping sparse models efficient when only a small part of the model runs for each token. For
-
Gemini API Cost Control for CI Agents Experimentation
By
–
This is great news for anyone who wants to run Gemini prompts in CI or let their agents know experiment with the Gemini API without fear of a nasty surprise bill https://t.co/y2q7TSsNU5
— Simon Willison (@simonw) 12 mars 2026This is great news for anyone who wants to run Gemini prompts in CI or let their agents know experiment with the Gemini API without fear of a nasty surprise bill
-

Grok 4.2 Achieves Major Rankings Improvement on BullshitBench
By
–
BullshitBench v2 Update: Grok 4.2 – massive jump in the rankings – 4.1 was ranked 54th and 72nd (out of 84) and now it took 13-16th spots. https://t.co/EdwOBStKa4 pic.twitter.com/6T6Tr0KsDr
— Peter Gostev (@petergostev) 12 mars 2026BullshitBench v2 Update: Grok 4.2 – massive jump in the rankings – 4.1 was ranked 54th and 72nd (out of 84) and now it took 13-16th spots.