"Kill vector databases" is a bold claim haha. We'll see that at scale! Curious what the alternative retrieval layer actually looks like in practice.
GENERATIVE AI
-

Claude Code Opus 4.6 1M Default Model and API Improvements
By
–
Tons of improvements shipped with this one:
– Opus 4.6 1M is now the default Opus model for Claude Code users on Max, Team, and Enterprise plans.
– No more long context price increase in the API.
– No beta header required in the API.
– Include up to 600 images in one request. -
Claude Reaches 1 Million Context Tokens in Production
By
–
https://claude.com/blog/1m-context-ga [Translated from EN to English]
→ View original post on X — @alexalbert__, 2026-03-13 18:22 UTC
-

NVIDIA Nemotron 3 Super Now Available in Perplexity Platform
By
–
NVIDIA’s Nemotron 3 Super is now available in Perplexity, Agent API, and Computer.
→ View original post on X — @perplexity_ai, 2026-03-13 18:15 UTC
-
Jensen Huang to Unveil AI and Computing Breakthroughs at GTC 2026
By
–
The countdown to #NVIDIAGTC is on. In just three days, our CEO Jensen Huang will take the stage at the SAP Center for the GTC 2026 keynote. Tune in Monday, March 16 at 11 a.m. PDT to be among the first to hear the newest breakthroughs in AI and accelerated computing. Get ready
-

JTok: Scaling LLMs with Token-Indexed Parameters
By
–
Can LLMs achieve massive capacity gains without massive increases in computational cost? YES, say researchers from Shanghai Jiao Tong University and Xiaohongshu! They introduce JTok, a novel scaling method that uses lightweight "token-indexed parameters" to intelligently
-
Real-time video captioning in browser with LFM2-VL WebGPU
By
–
Real-time video captioning in your browser, powered by LFM2-VL
— Maxime Labonne (@maximelabonne) 13 mars 2026
Another exceptional demo by the great @xenovacom! https://t.co/KRvDEpj138Real-time video captioning in your browser, powered by LFM2-VL Another exceptional demo by the great @xenovacom! Xenova (@xenovacom) Real-time video captioning in your browser with @LiquidAI's LFM2-VL model on WebGPU. Sending every frame to a server was never going to be the answer. Imagine the bandwidth, latency and cost. Local inference. No server costs. Infinitely scalable. This is the way. — https://nitter.net/xenovacom/status/2032504624024854673#m
→ View original post on X — @maximelabonne, 2026-03-13 17:46 UTC
-
Micro-benchmarks Don’t Measure True Reasoning Capabilities
By
–
These micro-benchmarks are fun but I've found the model that "wins" changes depending on the exact framing of the prompt. Pattern matching != reasoning.
-

Anthropic releases Opus 4.6 with 1M context window
By
–
Anthropic made Opus 4.6 with a 1M context window, generally available to all Claude Code users.
-
Lending stochastic parrots term, refining critique
By
–
Yeah, we’ve got to say it! (Well, after that, I did lend him the term “stochastic parrots,” which he doesn’t use, thanks @Fabien_Mikol
, I need to refine my critique)