Can you get a Mamba model to perform like a Transformer without adding Attention? Researchers from Apple, MILA, and Flat Iron Institute (including Abhinav Moudgil and Ningyuan Huang) have a breakthrough answer. They introduce a two-step distillation recipe: first, they convert
RESEARCH
-
User Wants More AI Content on Social Media Feed
By
–
im seeing so many irrelevant clips on x lately. I want my fy-page filled with AI stuff the way it was and not a second tiktok 🙁 this really makes me sad.
-

DeepSeek V4 Offers Cheapest Models Near Frontier Performance
By
–
More of my notes on DeepSeek V4 – the really big news is the pricing: both DeepSeek-V4-Flash and DeepSeek-V4-Pro are the cheapest models in their categories while benchmarking close to the frontier models from other providers https://
simonwillison.net/2026/Apr/24/de
epseek-v4/
… -

DeepSeek-V4 Breakthrough: Massive KV Cache Efficiency Gains
By
–
DeepSeek-V4 just dropped! And it's solving one AI's biggest problem today: It runs 1M-token context at 10% of the KV cache and 27% of the inference FLOPs of V3.2. Here's what that means. KV cache is the memory footprint your GPU holds for every token already in context. It
-
No malice assumed: TiKZ unicorns training not in OAI’s interest
By
–
I don’t see any reason to assume malice here—even if increased training on TiKZ unicorns is the true explanation, you’d expect it to happen naturally given the impact of Bubeck’s work. More generally it’s not in OAI’s interest to “bechmaxx” weird tests and they surely know this.
-

AI Model Self-Corrects Visual Grounding Errors With Confidence Scoring
By
–
What if a model could catch and correct its own mistakes while learning? Researchers from Peking University present a new AI method for visual grounding. Instead of just matching words to image regions, their system uses a "confidence score" to flag its unreliable guesses. It
-

Standard and Ultra High Context Efficiency in LLMs
By
–
1m Standard and ultra high context efficiency is what me excites me
-

DeepSeek Outshines OpenAI’s GPT-5.5 Release Strategy
By
–
Did Deepseek really wait until OpenAI released GPT-5.5 to steal the show?
-
New Open-Weight DeepSeek Model Announced with Good Benchmarks
By
–
And now a new DeepSeek model, and appears to be fully open weights. Good benchmarks, but with open models, that isn't always as meaningful. Should be live soon to actually try.
-

DeepSeek v4 Sets New SOTA Open Source Record
By
–
Deepseek v4 is a huge step upwards compared to DeepSeek 3, outperforms on SWE verified opus 4.6 and GPT-5.4 and sets a new record on Codeforces. Needs to be tested against opus 4.7 and GPT-5.5 tho and see if real world usage holds its promises. Big release! Sota open source
