The thing I most need a name for right now is that irresistible urge to announce to myself that AI labs are probably training on my benchmark as soon as I share an image of a pelican riding a bicycle anywhere on the internet.
RESEARCH
-

AI Wins Gold at Math Olympiad via Simple Unified Scaling
By
–
Cool! AI can win gold at the International Math Olympiad via Simple and Unified Scaling! Researchers from Shanghai AI Lab, CUHK, Tsinghua and PKU introduce SU-01. Their simple recipe: first train on proof-search and self-checking behaviors, then scale via two-stage
-
Opus 4.8 research workflow and GPT-5.5 feedback
By
–
Opus 4.8 formulated the hypotheses in advance, conducting data cleaning, did research on references, conducted analyses, did robustness checks, and put out the whole paper in LaTEX style. GPT-5.5 found one issue with a hallucinated result, and had other constructive feedback.
-

AI Models Understanding Chemical Principles
By
–
Building #AI models that understand chemical principles
by Anne Trafton @MIT Learn more: https://
bit.ly/4dtKrgb #ArtificialIntelligence #MachineLearning #ML -

AI agents wrote and reviewed an academic paper
By
–



I had Opus 4.8 in Claude Code write a sophisticated, if minor, academic paper from a archive of hundreds of de-identified research files from years ago I had to use GPT-5.5 Pro as a reviewer, it spotted one major error & some minor points. Opus corrected https://
embeddedness-gradient.netlify.app -
The Age of Async Agents: Devin’s Growth, AI Commits, and Cloud Engineering
By
–
🆕The Age of Async Agents: Devin’s 7x PR growth, 80% AI commits, background agents, memory, testing, & Open-Inspect https://t.co/x5Hw5S3egc@cognition cofounder + CPO @walden_yan and Open-Inspect creator @_colemurray explain why engineering is moving from local IDEs to cloud… pic.twitter.com/fciT77nJNI
— Latent.Space (@latentspacepod) 28 mai 2026The Age of Async Agents: Devin’s 7x PR growth, 80% AI commits, background agents, memory, testing, & Open-Inspect https://
latent.space/p/cognition @cognition cofounder + CPO @walden_yan and Open-Inspect creator @_colemurray explain why engineering is moving from local IDEs to cloud -
AI Model Comparison: Opus 4.8 vs. GPT-5.5 Trajectory
By
–
Opus 4.8 is clearly a strong model, but my impression is that Anthropic is increasingly playing catch-up with OpenAI rather than setting the pace. It feels like GPT-5.5 has shifted the benchmark again, and if OpenAI keeps this trajectory, GPT-5.6 could very plausibly become the
-
Decentralized Training Scales Video Models Without Data Centers
By
–
This is the part nobody is talking about yet.
— AI Highlight (@AIHighlight) 28 mai 2026
You can now train a serious video model without owning a single data center.
Bagel proved decentralized training scales from images to video, and the next stop is world models. https://t.co/vhPzDocy2BThis is the part nobody is talking about yet. You can now train a serious video model without owning a single data center. Bagel proved decentralized training scales from images to video, and the next stop is world models.
-
LocateAnything: Vision-Language Detection Model for AI Agents
By
–
This #CVPR2026 paper from our research team is trending #1 on @HuggingFace 🤗
— NVIDIA AI (@NVIDIAAI) 28 mai 2026
Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to… pic.twitter.com/2OGaQnUCnXThis #CVPR2026 paper from our research team is trending #1 on @HuggingFace Meet LocateAnything: a vision-language detection model that rethinks bounding box prediction. For AI agents and robots, “seeing” is only useful if a model can pinpoint where something is fast enough to
-

TaH: Skipping extra reasoning steps makes AI smarter
By
–
Your AI model is thinking too much, and skipping a few steps makes it smarter! Tsinghua University, Infinigence AI, and Shanghai Jiao Tong University introduce Think-at-Hard (TaH). Instead of looping every token through extra reasoning, TaH uses a lightweight decider to skip