DeepSWE is a new state-of-the-art open-source software engineering model trained entirely using reinforcement learning, based on Qwen3-32B. https://
together.ai/blog/deepswe Fantastic work from @togethercompute @Agentica_
@hardmaru
-

DeepSWE: New Open-Source Software Engineering Model with Reinforcement Learning
By
–
-
Pass@2 Scores Gap Below Pass@250 Levels Explained
By
–
The pass@2 scores are actually reported in the blog post, but as mentioned in the parent tweet, the gap is ~ 10-15% below pass@250 levels. What I meant in my previous tweet is to narrow this gap in the future.
-
Scaling LLM Inference Compute: Adaptive Branching Tree Search
By
–
Our paper: “Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search”
-
Field Advances Rapidly While ARC-AGI-2 Still Under Development
By
–
This field is moving too fast. We’re still working on ARC-AGI-2…
-

AB-MCTS Search Capability Evaluation with High Pass@k Metrics
By
–
Thanks @fchollet
. Indeed, the experiments used a large Pass@k which allowed us to focus on evaluating the search capability of AB-MCTS, rather than the official evaluation criteria based on k=2. We also used tasks in the public eval. Hopefully we’ll get down to Pass@2 someday! 🙂 -

Multi-LLM AB-MCTS Combination Outperforms on ARC-AGI-2 Benchmark
By
–
The Multi-LLM AB-MCTS combination of o4-mini + Gemini-2.5-Pro + DeepSeek-R1-0528, current frontier AI models, achieves strong performance on the ARC-AGI-2 benchmark, outperforming individual models by a large margin. Implementation of AB-MCTS on GitHub: https://
github.com/SakanaAI/treeq
uest
… -

Multiple LLMs Solve Previously Unsolvable ARC-AGI-2 Examples Together
By
–
Many ARC-AGI-2 examples that were unsolvable by any single LLM were solved by combining multiple LLMs. In some cases, an initially incorrect attempt by o4-mini is used by R1-0528 and Gemini-2.5-Pro as a hint to get to the correct solution. ARC-AGI-2 code: https://
github.com/SakanaAI/ab-mc
ts-arc2
… -
AB-MCTS: Inference-Time Scaling for Frontier AI Model Cooperation
By
–
Inference-Time Scaling and Collective Intelligence for Frontier AIhttps://t.co/3qSUaEixQU
— hardmaru (@hardmaru) 1 juillet 2025
We developed AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, achieving promising initial results on the ARC-AGI-2 benchmark.… pic.twitter.com/9guSv7Dtv0Inference-Time Scaling and Collective Intelligence for Frontier AI https://
sakana.ai/ab-mcts/ We developed AB-MCTS, a new inference-time scaling algorithm that enables multiple frontier AI models to cooperate, achieving promising initial results on the ARC-AGI-2 benchmark. -

Reading Technical Report with Attribution to Teknium
By
–
Reading their technical report (h/t @teknium
) https://
x.com/Teknium1/statu
s/1939518639323123917
… -
Schmidhuber Reflects on Surprising AI Research Results
By
–
“It’s nice work,” said Jürgen Schmidhuber (
@SchmidhuberAI
). “I think for many people, the results are surprising. Since I’ve been working on that topic for almost 40 years now, it’s maybe a little bit less surprising to me.”