If your benchmark relies on a static dataset or sampling from a static distribution densely known at training time, then it is fundamentally measuring memorization/retrieval. Which might be fine if you're looking for a retrieval benchmark! But don't confuse it with intelligence.
RESEARCH
-
AI diffusion model discovers Big Mac and creates optimized burgers
By
–
Finally, AI finds its ultimate uncontroversial use. A diffusion model trained on burger recipes "discovers the classic Big Mac without explicit supervision and generates novel burgers optimized for deliciousness, sustainability, or nutrition." ASI= automated slider intelligence
-

Stanford AI adds realistic personalities to train crisis workers
By
–
Stanford scholars developed a new way of adding human differences back into AI-generated text. With more realistic personalities, AI can simulate patients with specific symptom profiles for training crisis-line workers and clinicians. https://
hai.stanford.edu/news/todays-ai
-talks-like-nobody-new-research-gives-it-real-personality
… -
AGI requires innovative ideas, not just fast code writing
By
–
no, innovative ideas rendered into code will create AGI. we still need the ideas. (but yes writing code faster is helpful; then again the code they write tends not to be innovative)
-

ViQ: Text-Aligned Visual Quantized Representations at Any Resolution
By
–
ViQ Text-Aligned Visual Quantized Representations at Any Resolution
-
DanceOPD: On-Policy Generative Field Distillation
By
–
DanceOPD
— AK (@_akhaliq) 26 juin 2026
On-Policy Generative Field Distillation pic.twitter.com/8ZPiPVVgkEDanceOPD On-Policy Generative Field Distillation
-

PACE: Training optimizers for the averaged model you return
By
–
"Training for the Model You Return" Most LM pipelines return an EMA or averaged checkpoint, but optimizers still train like the final iterate is what matters. So if the output is an averaged model, can we shape training so that average gets better? This paper introduces PACE,
-

Autodata: An agentic data scientist to create high quality synthetic data
By
–
"Autodata: An agentic data scientist to create high quality synthetic data" If there's auto-research, shouldn't there also be a auto-data generation? In this new Meta paper, they proposed Autodata, which makes synthetic data generation work more like a data scientist, with an
-
Billions spent on AI, but memory was the real bottleneck
By
–
Do you understand the irony of what just happened?
— Robert Scoble (@Scobleizer) 26 juin 2026
We spent two years and BILLIONS of dollars making AI models smarter. What was actually holding them back was that they forget everything after each session.
It was memory. https://t.co/0krKlMh469Do you understand the irony of what just happened? We spent two years and BILLIONS of dollars making AI models smarter. What was actually holding them back was that they forget everything after each session. It was memory.
-
Anthropic advances study of Claude’s economic impact with hourly data
By
–
To keep pace with AI progress, we're advancing how we study Claude's economic impact. Hourly sampling and survey data show us how the cadences of life shape usage, what people produce with Claude, and how perceptions of AI's impact may be changing.