was running some evals this weekend and claude kept trying to get me to go to bed
LLMS
-
Long context: expensive fake memory, not true recall
By
–
Long context was never memory. It was a very expensive way to pretend your agent remembered something it was just rereading badly.
-
Agents lack session memory despite context window focus, HydraDB proposes new layer
By
–
For two years the whole conversation was about context window size.
— Chubby♨️ (@kimmonismus) 1 juin 2026
Meanwhile the actual problem never moved: agents don't remember anything between sessions. We kept patching it with RAG and manual context injection and calling that memory.
HydraDB is going at the layer… https://t.co/st09X5C45VFor two years the whole conversation was about context window size. Meanwhile the actual problem never moved: agents don't remember anything between sessions. We kept patching it with RAG and manual context injection and calling that memory. HydraDB is going at the layer
-

ContinuousBench: Hard Leakage-Proof DP Synthetic Text Benchmark
By
–
Does DP synth text transfer useful knowledge or just superficial style mimicking? Existing benchmarks: saturated Introducing ContinuousBench: a hard (curr methods fail at ε=100! ) & leakage-proof benchmark for DP synth text! Followup to our #ICML2024 best paper 1/n
-
Benchmarking confirms reasoning-only model requires thinking mode
By
–
Awesome, thanks for benchmarking it! It makes sense, the model is reasoning-only so thinking mode is mandatory here.
-

Claude Code deletes session traces after a month
By
–
i was today years old when i learned that claude code deletes your session traces after a month
-
Cosmos 3 Nano and Super models on Hugging Face with datasets
By
–
6/ Two sizes, both live on Hugging Face right now: Cosmos 3 Nano (8B) — runs on a single workstation GPU for real-time robotics
Cosmos 3 Super (32B) — datacenter-grade, max quality Plus six open datasets and full post-training scripts on GitHub. -

MiniMax M3: Open weights, 1M tokens, native multimodal AI
By
–
MiniMax M3 just raised the bar—open weights, 1M-token context, and native multimodal from day one. Visit http://
futurepedia.io, the leading AI tools directory. -

GrepSeek: Training Search Agents for Direct Corpus Interaction
By
–
GrepSeek Training Search Agents for Direct Corpus Interaction
-
NVIDIA announces Nemotron 3 Ultra, 550B open-weight model
By
–
NVIDIA announced an upcoming release of Nemotron 3 Ultra later this week, a 550B-parameter open-weight model.
— 🚨 AI News | TestingCatalog (@testingcatalog) 1 juin 2026
According to Artificial Analysis, it is positioned as the most intelligent open-weight model from the US lab.
Soon 👀 https://t.co/rSWH9n9oIh pic.twitter.com/8XiWwDRwJzNVIDIA announced an upcoming release of Nemotron 3 Ultra later this week, a 550B-parameter open-weight model. According to Artificial Analysis, it is positioned as the most intelligent open-weight model from the US lab. Soon
