Can we finally trust what LLMs are really thinking? Yutong Gao and researchers from Peking University, Purdue, and Nanjing University present a new survey on intrinsic interpretability — building transparency directly into model architecture rather than relying on post-hoc
RESEARCH
-
Improving Claude’s Safe Behavior via Training and Response Rewriting
By
–
We experimented with training Claude on examples of safe behavior in scenarios like our evaluation. This had only a small effect, despite being similar to our evaluation. We got further by rewriting the responses to portray admirable reasons for acting safely.
-
Anthropic Research Eliminates Harmful Behavior in Claude 4
By
–
New Anthropic research: Teaching Claude why. Last year we reported that, under certain experimental conditions, Claude 4 would blackmail users. Since then, we’ve completely eliminated this behavior. How?
-

New ICML26 Paper on Sparse Transformer Kernels for NVIDIA GPUs
By
–
Great collab with @SakanaAILabs on an #ICML26 paper about sparse transformer kernels + formats optimized for modern NVIDIA GPU execution. • TwELL sparse packing
• Fused CUDA kernels
• 20%+ inference/training speedups at scale Paper + code below -

Representational Geometry Influences Uncertainty in Large Language Models
By
–
"Representational Curvature Modulates Behavioral Uncertainty in LLMs" Most interpretability work studies what features live inside LLMs. But this paper studies something deeper, which is the geometry of the whole evolving representation. The key result is that representational
-
Optimiser les LLM avec la sparsité adaptée au GPU
By
–
How do we make LLMs faster and lighter? Don’t force the GPU to adapt to sparsity. Reshape the sparsity to fit the GPU! ⚡️
— Sakana AI (@SakanaAILabs) 8 mai 2026
Excited to share our new #ICML2026 paper in collaboration with @NVIDIA: "Sparser, Faster, Lighter Transformer Language Models". This work introduces new… pic.twitter.com/ehByWHIh6IHow do we make LLMs faster and lighter? Don’t force the GPU to adapt to sparsity. Reshape the sparsity to fit the GPU! Excited to share our new #ICML2026 paper in collaboration with @NVIDIA
: "Sparser, Faster, Lighter Transformer Language Models". This work introduces new -

Continuous-Time Distribution Matching for Few-Step Diffusion Distillation
By
–
Continuous-Time Distribution Matching for Few-Step Diffusion Distillation paper: https://
huggingface.co/papers/2605.06
376
… -

Apple introduces TIDE: Every Layer Knows the Token Beneath the Context
By
–
Apple presents TIDE Every Layer Knows the Token Beneath the Context paper: https://
huggingface.co/papers/2605.06
216
… -
LLMs’ Math Progress Accelerates Rapidly
By
–
It's surprising to see the pace of progress of LLMs (which until not long ago were very bad at math) achieving ever greater milestones. Take a look at the graph above. The progress curve is getting faster and faster. I'll leave the link to the paper here: https://
arxiv.org/abs/2605.06651 -

MARBLE: Multi-Aspect Reward Balance for Diffusion RL
By
–
MARBLE Multi-Aspect Reward Balance for Diffusion RL paper: https://
huggingface.co/papers/2605.06
507
…
