Not Even Bronze: Evaluating LLMs on 2025 International Math Olympiad https://
matharena.ai/imo/ Nice blog post from the team behind MathArena: Evaluating LLMs on Uncontaminated Math Competitions (
https://
arxiv.org/abs/2505.23281) providing independent analysis of LLM performance on IMO.
@hardmaru
-

LLMs Performance on 2025 International Math Olympiad Evaluation
By
–
-

Paradigm Shifts in AI: Collective Intelligence and Artificial Life
By
–
Blaise Agüera explains interrelated paradigm shifts which he believes are core to the future development in AI. I like his take on collective intelligence, the future of artificial life research, and the (somewhat philosophical) discussions about consciousness and theory of mind.
-
AI Evolution Reshaping Our Understanding of Intelligence
By
–
New Essay by @BlaiseAguera (
@Google
): “AI Is Evolving — And Changing Our Understanding Of Intelligence” Advances in AI are making us reconsider what intelligence is and giving us clues to unlocking AI’s full potential. -
Foundation Models for NP-Hard Optimization Problems
By
–
See our recent work on applying foundation models to tackle challenging NP-Hard optimization problems:https://t.co/uX2YFnqsOYhttps://t.co/MDPwf8YTkx
— hardmaru (@hardmaru) 18 juillet 2025See our recent work on applying foundation models to tackle challenging NP-Hard optimization problems: https://
sakana.ai/ale-bench -

FakePsyho Wins AtCoder World Tour Finals 2025 Heuristic
By
–
Congrats to @FakePsyho for winning AtCoder World Tour Finals 2025 Heuristic Humanity has prevailed (for now!) Thanks OpenAI for sponsoring #AWTF2025, and getting #2 on this grand challenge. Proud of @SakanaAILabs & @AtCoder
’s ALE-Agent for reaching #5, on a shoestring budget! -

World Models Research Figure Still Used Seven Years Later
By
–
Nice to see a figure from our 2018 paper about World Models is still being used in 2025. https://
worldmodels.github.io https://
arxiv.org/abs/1803.10122 -
Video on Fractured Entangled Representation Hypothesis in Deep Learning
By
–
Nice video discussing the recent thought-provoking paper: “Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis” https://t.co/Wu17RXpUaV
— hardmaru (@hardmaru) 5 juillet 2025
By @akarshkumar0101 @jeffclune @joelbot3000 @kenneth0stanley
15min @MLStreetTalk video↓ https://t.co/H9mCCEQpb5Nice video discussing the recent thought-provoking paper: “Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis” https://
arxiv.org/abs/2505.11581 By @akarshkumar0101 @jeffclune @joelbot3000 @kenneth0stanley 15min @MLStreetTalk video↓ -
Soham Parekh couldn’t pass Sakana AI’s tough hiring process
By
–
Someone like Soham Parekh would never be able to pass Sakana AI’s hardcore job application process.
-
Top open-source software engineering models dominated by Qwen and DeepSeek variants
By
–
The top open SWE models are all Qwen3, Qwen2, QwQ, Deepseek-V3 or R1-based.