And here we are! Finally, after nearly 15 months of waiting. These models are monsters though. Even the “small” one will require substantial hardware to run on.
LLMS
-

DeepSeek-V4 Introduces Advanced Attention Techniques for Million Token Context
By
–
"DeepSeek-V4 Technical Report" A 58 page paper with brand new attention techniques: Heavily Compressed Attention (HCA) & Compressed Sparse Attention (CSA). This hybrid attention setup enables V4 to hit 1 million context. DeepSeek-V4-Pro is now the largest OS model ever, with
-
DeepSeek V4 Paper Release Announcement
By
–
Check out the full paper here! https://
alphaxiv.org/abs/deepseek-v4 -
Sakana AI Launches Fugu Multi-Agent Orchestration System Beta
By
–
Sakana AI is launching the beta test of "Sakana Fugu," a new commercial AI product—a multi-agent orchestration system Blog: https://
sakana.ai/fugu-beta/#Jap
anese
… This is a system that dynamically coordinates multiple frontier foundation models, autonomously selecting the optimal -

DeepSeek Mocks Anthropic’s Claude in AI Competition
By
–
Even DeepSeek is now making fun of Anthropic's Claude.
-

Cross-Architecture Distillation Recipe for Mamba Models
By
–
Attention to Mamba: A Recipe for Cross-Architecture Distillation Paper: https://
arxiv.org/abs/2604.14191 -

Mamba Models Match Transformer Performance Without Attention
By
–
Can you get a Mamba model to perform like a Transformer without adding Attention? Researchers from Apple, MILA, and Flat Iron Institute (including Abhinav Moudgil and Ningyuan Huang) have a breakthrough answer. They introduce a two-step distillation recipe: first, they convert
-
QClaw International Beta Now Free with 40M Daily Tokens
By
–
4/ Their international beta is now accessible and completely free if anyone wants to test it out. Capacity: 40M tokens/day. Relationship with OpenClaw: QClaw is built on the OpenClaw open-source framework with an end-to-end consumer encapsulation layer. All QClaw-developed
-
DeepSeek v4 Performance Comparison and Visual Generation Capabilities
By
–
If you want to see what DeepSeek v4 is like, there's no better way than looking at my face alongside the dozens of generations I did side by side with other models
-
Architectural Principles for Building AI Agents and Automations
By
–
What this means practically for anyone building agents, prompts, or automations right now: – Stop treating memory as a storage problem – Stop renting your agent's intelligence from the labs – Instrument outcomes, not just inputs – Every interaction should be a labeled
