I never complain! The peak was when we had 4o and o4 at the same time
LLMS
-

Apple Presents Stochastic KV Routing for Adaptive Cache Sharing
By
–
Apple presents Stochastic KV Routing Enabling Adaptive Depth-Wise Cache Sharing paper: https://
huggingface.co/papers/2604.22
782
… -
Nemotron 3 Nano Omni: NVIDIA’s Fully Open Source AI Model
By
–
Built on NVIDIA’s open ecosystem, Nemotron 3 Nano Omni is fully open source, including:
• Open weights
• Open data
• Open recipes Read the blog for more details -

Nemotron 3 Nano Omni: Unified Multimodal Architecture for AI Subagents
By
–
Nemotron 3 Nano Omni was designed for powering subagents. Instead of stitching together separate models for language, vision, and speech, it ties them into a single architecture that more efficiently feeds context to orchestrators.
-

Nemotron 3 Nano Omni: Efficient Open Multimodal Model Released
By
–
Meet Nemotron 3 Nano Omni 👋
— NVIDIA AI (@NVIDIAAI) 28 avril 2026
Our latest addition to the Nemotron family is the highest efficiency, open multimodal model with leading accuracy.
30B parameters. 256K context length. 🧵👇 pic.twitter.com/j4SPpU9SaIMeet Nemotron 3 Nano Omni Our latest addition to the Nemotron family is the highest efficiency, open multimodal model with leading accuracy. 30B parameters. 256K context length.
-

New Method to Train and Evaluate AI Agents Efficiently
By
–
Vibe train your AI agents. There's a new method that could replace LLM-as-a-judge for production agents. Most teams rely on a giant LLM as a judge to evaluate and guard their agent. But it has two major drawbacks: – It's slow and expensive at inference time
– It often misses -

24GB RAM Model Achieves 150K Token Context Window
By
–
24gb ram 150k context in: about 200
out: screenshot below ↓ -
Live Session on Building LLM Knowledge Bases Tomorrow
By
–
Live session on how to build LLM knowledge bases tomorrow: https://
academy.dair.ai/dashboard/even
ts/cmnivyzyp001n04k1rnozju2n
… -

Poolside Releases Laguna XS and M Coder Models on Hugging Face
By
–
If you've ever wondered what Poolside was up to… They've been cooking!
They just released Laguna XS and M, their first ever public models, on Hugging Face:
– Two coder models: one is 225B params, of which 23B active, the other is 33B-3A.
– hybrid attention: global v. sliding -

Claude Detects When Users Switch to Codex LLM
By
–
Someone discovered that Claude knows when you've been cheating on it with Codex. Lucas maintains an open relationship with his LLMs. Source: https://
reddit.com/r/ClaudeAI/com
ments/1sxe46v/claude_knows_when_you_cheat_on_it_with_codex/
…