The response to Open Benchmarks Grants has been strong. We’re seeing rigorous proposals across open datasets, benchmarks, and evaluation methods — exactly the kind of infrastructure AI needs. Our research teams have been energized reviewing the first wave. First review
RESEARCH
-
MIT CSAIL Research Update April 2026
By
–
This is MIT CSAIL: https://t.co/SIrlbVeuiV pic.twitter.com/Ph2yJHG22L
— MIT CSAIL (@MIT_CSAIL) 25 février 2026This is MIT CSAIL: https://
bit.ly/4cIwpak -
AI-Powered Materials Science: New Era of Automated Discovery
By
–
🔬 New Science pod with @cusp_ai!
— Latent.Space (@latentspacepod) 25 février 2026
We are entering a new era where materials science and discovery is transitioning from slow, manual experimentation, to a high-speed search problem powered by generative AI and "physics processing units." @wellingmax argues that the foundation… pic.twitter.com/UKZ5xH9NK4🔬 New Science pod with @cusp_ai! We are entering a new era where materials science and discovery is transitioning from slow, manual experimentation, to a high-speed search problem powered by generative AI and "physics processing units." @wellingmax argues that the foundation of all modern technology—from GPUs to climate solutions—is a materials problem, and that unifying the mathematics of stochastic thermodynamics with generative AI will unlock a new paradigm of automated scientific discovery.
→ View original post on X — @wellingmax, 2026-02-25 17:50 UTC
-

The Dragon Book: Four Decades of Compiler Design Legacy
By
–
How you know you're getting older: the "dragon book" was published nearly forty years ago. https://
shorturl.at/DswlP -

AI Breakthrough: Aletheia Solves Mathematical and Scientific Discovery Problems
By
–
For more information about Aletheia, powered by Gemini #DeepThink, and our works in AI for mathematical and scientific discovery at @GoogleDeepMind and @GoogleResearch, check out our announcement last week nitter.net/lmthang/status/2021631…! Thang Luong (@lmthang) 6 months in, after the IMO-gold achievement, I’m very excited to share another important milestone: AI can help accelerate knowledge discovery in mathematics, physics, and computer science! We’re sharing Two new papers from @GoogleDeepMind and @GoogleResearch that explore how Gemini #DeepThink together with agentic workflows can empower mathematicians and scientists to tackle professional research problems. Some highlights: The first paper built a research agent #Aletheia, powered by an advanced version of Gemini Deep Think, that can autonomously produce publishable math research and crack open Erdős problems. The second paper, built on similar agentic reasoning ideas, helped resolve bottlenecks in 18 research problems, across algorithms, ML and combinatorial optimization, information theory and economics. See the thread for details about the two papers and the joint blog post. — https://nitter.net/lmthang/status/2021631397614731563#m
-

FirstProof Problem 7 Solution Confirmed by Original Mathematician
By
–
The correctness of our solution to FirstProof problem 7 is also confirmed by Jim Fowler, the mathematician who conjectured the question originally! See github.com/google-deepmind/s… for all our transcripts and solutions (both correct and incorrect ones!) as well as public discussion of P7 at icarm.zulipchat.com/#narrow/….
-

Aletheia AI Solves Open Math Problem P7 Successfully
By
–
This is a remarkable milestone in which our agent can work on a research problem for a very long time, then come back and tell us if it has succeeded or failed! We visualize the inference cost Aletheia decided to spend on each candidate solution (as a multiple of the inference cost of for solving Erdős-1051, see our previous work nitter.net/lmthang/status/2018354…). P7 is extremely interesting. It has been an open problem for several years, and nobody else came close to solving it in the FirstProof contest per @tonylfeng. We initially thought Aletheia had no chance; turned out it was right! Aletheia spent most compute on P7, 16x amount we used for Erdős-1051. Remarkably, per @kimshmath, "This was the first case that I have ever seen that an AI applies several deep mathematical results (by Cartan/Leray/Borel/Atiyah/Quillen/Novikov/Kasparov…) flawlessly. It is a very unique instance."
-

Aletheia Agent Solves 6 of 10 FirstProof Math Challenge Problems
By
–
Exciting results in AI math research! We use Aletheia agent, powered by Gemini 3 Deep Think, to tackle the FirstProof challenge. Operating completely autonomously, Aletheia successfully solved 6 out of the 10 problems. Check out the full paper for details on the methodology and expert evaluations. arxiv.org/abs/2602.21201
-
Test-Time Training with KV Binding as Linear Attention
By
–
Test-Time Training with KV Binding Is Secretly Linear Attention
-

Aletheia solves 6 of 10 FirstProof problems using Gemini DeepThink
By
–
We ran two Aletheia versions (differing only by base model) powered by Gemini #DeepThink. Together, they solved 6/10 problems (2, 5, 7, 8, 9, 10) per majority expert assessments. Full transparency on our FirstProof interpretation and experiments: arxiv.org/abs/2602.21201. Evaluation is extremely hard! Only a handful of experts can even understand these problems. As such, we have conducted our study very carefully! Crucially, our solutions were generated without any human intervention and submitted within the timeframe of the FirstProof challenge. The lead author of FirstProof confirmed that fact in the public Zulip discussion of our solutions icarm.zulipchat.com/#narrow/….
