This is still 0.5T, but a more recent training checkpoint. 1T model is ~5 days away from finishing initial training. Will be a major step change improvement in coding, long context and skills. The SpaceXAI model factory is finally working. Should be an improved base model
LLMS
-

AI Model Distillation: How Stronger Models Train Weaker Ones
By
–
Everyone is accusing everyone of “stealing AI”
But almost nobody is explaining what’s actually happening. Distillation. → Query a stronger model at scale
→ Collect outputs (reasoning, code, decisions)
→ Train your own to imitate it No weights. Just behavior. This worked in -
Grok 4.3 Beta Release Improvements Daily Updates
By
–
Grok 4.3 is still an early beta that will improve almost every day, but try it out!
— Elon Musk (@elonmusk) 17 avril 2026
We will publish release notes as we fix bugs and add functionality. https://t.co/s64QqC5tSsGrok 4.3 is still an early beta that will improve almost every day, but try it out! We will publish release notes as we fix bugs and add functionality.
-
Experienced developers can now leverage Claude and GPT models
By
–
Lisez bien. Sur la base de ce que je peux constater et que j'expérimente quotidiennement, j'affirme que tout développeur informatique expérimenté qui utilise @AnthropicAI #Claude Code et son modèle Opus 4.7 et également @OpenAI #Codex et son modèle GPT-5.4 peut désormais
-

Lennybot AI chatbot launches natively on Substack platform
By
–
New: Lennybot now lives natively within my Substack. Trained on all of my ~350 newsletter posts and ~300 podcast interviews. Check it out, ask it a question: https://
lennysnewsletter.com/cc/lenny-bot You can also text or call him anytime: +1 (877) 537-9455 -

RDU Architecture Solves Token Generation Decode Bottleneck
By
–
Decode is the bottleneck you actually feel. The RDU attacks it differently. 🦾
— SambaNova (@SambaNovaAI) 17 avril 2026
Data streams to compute, not the other way around.
Three-tier memory. PCU/PMU grids. No kernel-by-kernel stalling. That's fast token generation at the architecture level.
🔗 Learn more:… pic.twitter.com/yy2qqkquaVDecode is the bottleneck you actually feel. The RDU attacks it differently. Data streams to compute, not the other way around. Three-tier memory. PCU/PMU grids. No kernel-by-kernel stalling. That's fast token generation at the architecture level. Learn more:
-
Claude AI Models Show Different Help Capabilities: Opus 4.7 Explained
By
–
Every @claudeai model has a different idea of how much to help.@alexalbert__ from Anthropic explains why—and what to expect from Opus 4.7. pic.twitter.com/ByiHOdJZf1
— Dan Shipper 📧 (@danshipper) 17 avril 2026Every @claudeai model has a different idea of how much to help. @alexalbert__ from Anthropic explains why—and what to expect from Opus 4.7.
-
Claude Code grep capability transforms codebase context limitations
By
–
For me it was when Claude Code got good at grep, since then it didn't matter if the whole codebase fitted in the context or not
-
Opus 4.7 Performance Breakdown Across Coding Writing Spreadsheets
By
–
here is our full Opus 4.7 vibe check we broke down how it performs on coding, writing, spreadsheets and more: