Can you get a Mamba model to perform like a Transformer without adding Attention? Researchers from Apple, MILA, and Flat Iron Institute (including Abhinav Moudgil and Ningyuan Huang) have a breakthrough answer. They introduce a two-step distillation recipe: first, they convert
MACHINE LEARNING
-
User Wants More AI Content on Social Media Feed
By
–
im seeing so many irrelevant clips on x lately. I want my fy-page filled with AI stuff the way it was and not a second tiktok 🙁 this really makes me sad.
-
Architectural Principles for Building AI Agents and Automations
By
–
What this means practically for anyone building agents, prompts, or automations right now: – Stop treating memory as a storage problem – Stop renting your agent's intelligence from the labs – Instrument outcomes, not just inputs – Every interaction should be a labeled
-
The Architectural Trade-off Between Retrieval and Learning Systems
By
–
The moat math is brutal once you see it. A retrieval system plateaus at the quality of its extractor. As good on day 1,000 as day 1. And most of what it "knows" belongs to whichever lab hosts it. A learning system you control compounds every interaction into a corpus of (state
-
Neuro-inspired Three-tier Memory Architecture for AI Agents
By
–
The clearest frame I've seen on this comes from these guys @midbrain_ai. Their point: the brain doesn't run one memory system. It runs three. – Episodic: what happened (raw traces) – Semantic: what it means (abstracted patterns) – Procedural: how I now behave (baked into the
-
Economic trade-offs in AI agent memory: retrieval vs. procedural
By
–
Layer 3 is where the economics flip. Retrieval memory costs tokens on every call. The preference gets looked up, injected, re-reasoned — every time. Procedural memory costs zero tokens at inference. The behavior lives in the weights. The agent just acts correctly. At scale,
-

DeepSeek V4 Offers Cheapest Models Near Frontier Performance
By
–
More of my notes on DeepSeek V4 – the really big news is the pricing: both DeepSeek-V4-Flash and DeepSeek-V4-Pro are the cheapest models in their categories while benchmarking close to the frontier models from other providers https://
simonwillison.net/2026/Apr/24/de
epseek-v4/
… -

Indian AI Startups: Proven Solutions for Regulatory Automation
By
–
If you’re building AI solutions, this is where you prove it. We’re looking for Indian companies & DPIIT-recognised startups, teams with proven, proprietary AI solutions and builders with real-world deployment experience. Your solution could help in automating regulatory
-

DeepSeek-V4 Breakthrough: Massive KV Cache Efficiency Gains
By
–
DeepSeek-V4 just dropped! And it's solving one AI's biggest problem today: It runs 1M-token context at 10% of the KV cache and 27% of the inference FLOPs of V3.2. Here's what that means. KV cache is the memory footprint your GPU holds for every token already in context. It
-
No malice assumed: TiKZ unicorns training not in OAI’s interest
By
–
I don’t see any reason to assume malice here—even if increased training on TiKZ unicorns is the true explanation, you’d expect it to happen naturally given the impact of Bubeck’s work. More generally it’s not in OAI’s interest to “bechmaxx” weird tests and they surely know this.
