Last but not least, we experimented with a Gemini-based language reasoning system that showed great promise at this year’s IMO problems. This system doesn’t require the problems to be translated into a formal language and can directly generate and verify solutions in
@lmthang
-

DeepMind Releases AlphaGeometry 2 with Major AI Innovations
By
–
We created #AlphaGeometry 2 with lots of innovations: (a) one order of magnitude more synthetic data for the language model (Gemini-based, trained from scratch); (b) a symbolic engine that are two orders of magnitude faster, and (c) a novel knowledge-sharing mechanism to tackle
-

Neuro-symbolic AI: Language Models and Symbolic Reasoning
By
–
Neuro-symbolic paradigm strikes again! If in #AlphaGeometry, the creative language model (System 1) suggests insights for the reliable symbolic engine (System 2) to complete a proof, we see that pattern again in #AlphaProof. The language model suggests key proof steps in a
-
AI Solves 4 IMO Problems Despite Adversarial Combinatorics Setup
By
–
This year IMO problems were extremely adversarial to us with only 1 geometry, but 2 combinatorics (out of 6). It is also very difficult for humans: Terrence Tao admitted at the closing ceremony that P5 combinatorics screwed him! Despite so, we were able to solve 4 problems (P1/P6
-
AI Reaches Silver Medal Level in Mathematics at IMO 2024
By
–
Super thrilled to share that our AI has has now reached silver medalist level in Math at #imo2024 (1 point away from 🥇)! Since Jan, we now not only have a much stronger version of #AlphaGeometry, but also an entirely new system called #AlphaProof, capable of solving many more… pic.twitter.com/EyoUGrTALD
— Thang Luong (@lmthang) 25 juillet 2024Super thrilled to share that our AI has has now reached silver medalist level in Math at #imo2024 (1 point away from )! Since Jan, we now not only have a much stronger version of #AlphaGeometry, but also an entirely new system called #AlphaProof, capable of solving many more
-
Data Decontamination and Generalization in Advanced Math Benchmarks
By
–
That's a valid point. The team tried hard to decontaminate the data. Also note that there's generalization into more challenging benchmarks such as HiddenMath and IMO-Bench.
-
Gemini 1.5 Pro Achieves 90% MATH Benchmark Milestone
By
–
Also astonishing to see how fast the field has advanced: Gemini 1.5 Pro was first to surpass the 90% mark on MATH (6.9% 3 years ago). For comparison, on ImageNet, it took us close to 10 years to achieve the same from AlexNet (40%) to Meta Pseudo Labels (90.2%, our work
-

IMO-Bench: Gemini Math Achieves 25% on Mathematical Olympiad Problems
By
–
Glad to see us keep pushing the frontier on maths! Also a bit of a teaser to IMO-Bench (after #alphageometry), which my team (with @quocleix
's support) built & the best model only got 25% 🙂 Check out the cool answer of the Gemini Math on an APMO problem in @OriolVinyalsML
's -

Major AI Breakthroughs: RT-2, SIMA, GNoME, and AlphaFold3
By
–
Together with other cool works such as RT-2, SIMA, GNoME, and certainly AlphaFold3! https://t.co/Hgn6NVJMDC
— Thang Luong (@lmthang) 15 mai 2024Together with other cool works such as RT-2, SIMA, GNoME, and certainly AlphaFold3!
-

AlphaGeometry Highlighted at Google I/O Keynote
By
–
Just realized that #AlphaGeometry was highlighted at the beginning of @demishassabis
's keynote talk at #GoogleIO today. I was watching the keynote but somehow missed it, until @quocleix told me so :D. Stay tune for more updates from our team later in the year, cc @thtrieu_
!
