“Attention Drift: What Autoregressive Speculative Decoding Models Learn” Speculative decoding makes LLM inference faster, but drafters break under small template changes and long context. But why? This paper shows that as the drafter predicts more tokens, its attention drifts
MACHINE LEARNING
-

Improving LLM Embedding Representations via Mean-Pooling of Generated Tokens
By
–
“The Truth Lies Somewhere in the Middle of the Generated Tokens” LLMs don’t store the meaning of a prompt in one hidden state. As they generate, the meaning gets spread across many token embeddings. So this paper propose a mean-pool over generated token embeddings instead of
-
Opus 4.7 2.5x speed at 6x cost surprises user
By
–
i thought this was a joke. Opus 4.7 2.5x speed at 6x cost. what
-

Subquadratic Unveils New SubQ Model Using Subquadratic Sparse Attention
By
–
What if LLMs could read a million tokens without exploding in cost? Subquadratic, an AI research startup, just unveiled SubQ — the first model built on fully subquadratic sparse attention (SSA). Instead of comparing every token pair, SSA routes attention only to the truly
-
AI Agents Workflow: Codex, Claude Code, and Hermes Orchestration
By
–
Codex /goal builds it.
— Shubham Saboo (@Saboo_Shubham_) 12 mai 2026
Claude Code /goal review and refines it.
Hermes /goal manages the orchestration and handoff.
All tracked on a single Kanban Board and agents keep running in the loop. pic.twitter.com/WAIr8zCP4oCodex /goal builds it. Claude Code /goal review and refines it. Hermes /goal manages the orchestration and handoff. All tracked on a single Kanban Board and agents keep running in the loop.
-
LLMs Memorization vs Overfitting Clarification
By
–
Memorization does not imply overfitting. Overfitting is strictly about what happens on non-training data. So, e.g., just because LLMs memorize data doesn’t make them stochastic parrots. What matters is what they do with it, and they typically paraphrase it quite appropriately.
-
DeepMind’s AI pointer understands context and responds to voice
By
–
Google DeepMind just reinvented the mouse pointer.
— Chubby♨️ (@kimmonismus) 12 mai 2026
Since Doug Engelbart's demo in 1968, the little arrow on your screen has barely changed. Until now.
The new AI pointer sees what you're pointing at, understands the context, and responds to your voice. You point at an image of… https://t.co/z9cNgtMvOMGoogle DeepMind just reinvented the mouse pointer. Since Doug Engelbart's demo in 1968, the little arrow on your screen has barely changed. Until now. The new AI pointer sees what you're pointing at, understands the context, and responds to your voice. You point at an image of
-

NVIDIA Earth-2 and PhysicsNeMo Accelerate Severe Weather Prediction
By
–
Discover how @ColoradoStateU is revolutionizing severe weather prediction by using NVIDIA Earth-2 and PhysicsNeMo to extend hailstorm lead times from minutes to hours. This collaboration pairs generative AI with high-resolution radar data to deliver real-time, scalable
-
AI Interfaces Evolving Beyond Chat Towards Context-Aware Agents
By
–
“AI needs more than chat” is probably the most important line here. Chat interfaces were never the end product. They were just the temporary UI until agents got real context.
