very cool! you should store all the intermediary checkpoints on Hugging Face for permanence (and community visibility/sharing):
CODE
-
Codex Spark Best Practices Guide and Comprehensive Tips
By
–
for more comprehensive tips, check out our guide to Codex Spark: cerebras.ai/blog/codex-spark…
-
Container Issues with Nvidia Docker and Spark Setup
By
–
Maybe a container issue. I know Nvidia recommends their docker containers, but I am usually just running it on the Spark directly.
-
Cerebras Codex Spark Revolutionizes Coding with 1,200 Tokens Per Second
By
–
Our coding workflows were designed to accommodate slow inference. @OpenAI's Codex Spark powered by @cerebras changes the game.
— Cerebras (@cerebras) 12 mars 2026
Here's how we make the most out of 1,200 tokens per second, with @MilksandMatcha. pic.twitter.com/vv4a80wfFAOur coding workflows were designed to accommodate slow inference. @OpenAI's Codex Spark powered by @cerebras changes the game. Here's how we make the most out of 1,200 tokens per second, with @MilksandMatcha.
-
Agents Store Intermediary Checkpoints on Hugging Face
By
–
your agents could store the intermediary checkpoints on HF for persistence (+ sharing of course). we just released new storing options:
-
Working fine with Ollama and llama.cpp for Nemotron model
By
–
Weird, what issue are you having. Works for me both in ollama and native llama.cpp I am using unsloth/NVIDIA-Nemotron-3-Super-120B-A12B-GGUF:MXFP4_MOE
-
Building Large Language Models From Scratch YouTube Series
By
–
I actually have a Build an LLM From Scratch YouTube series
-

10 GitHub Repos to Master AI Engineering in 2026
By
–
If you want to become an AI engineer in 2026, spend less time collecting tutorials and more time reading real code, real notebooks, and real systems. Here are 10 GitHub repositories that can teach you more practical AI engineering than many expensive courses: 1) AI Agents for
-

Scalable MoE Training Efficiency with Megatron Core
By
–
“Scalable Training of Mixture-of-Experts Models with Megatron Core” This NVIDIA MoE report walks through the hard part of MoE training. The key is not to add more parameters, but keeping sparse models efficient when only a small part of the model runs for each token. For