Top AI Papers you Can't miss from last Week: We summarized everything from RL vs. SFT training, Hallucination Mitigation with Multi-agent systems, Microsoft's New FP4 Quantization, and more. Here's everything you need to know:
LLMS
-
Agentic Performance and GAIA Benchmark Analysis
By
–
With an agentic setup, quite obvious that scores would increase! The comparison against other agents on the GAIA benchmark is more interesting.
-
Appending Wait to Model Generation as Path to AGI
By
–
« appending "Wait" multiple times to the model's generation » => current most likely path to AGI
-
Using o1 as Reward Model for Output Selection
By
–
The code here yes, but concept is likely. It is quite likely they used o1 as a reward model/critic to choose from a group of outputs
-

DeepSeek-R1 and o1 Performance on Real-World Tasks
By
–
Beyond benchmarks: How DeepSeek-R1 and o1 perform on real-world tasks
by @BenDee983 @VentureBeat Read more: https://
buff.ly/3WJvca9 #ArtificialIntelligence #MI #MachineLearning cc: @karpathy @alvinfoo @ogrisel -
LLMs as Mainframe Computers: Early Stage AI Development
By
–
“Today’s LLMs are the ‘Mainframe Computers’ of our generation”
— hardmaru (@hardmaru) 3 février 2025
I was on @BloombergTV today discussing @SakanaAILabs, and to share my view that today’s LLMs are our generation’s “Mainframe Computers”. We are still in the very early stages of AI, and it is inevitable, due to… pic.twitter.com/PHJNL1GdCM“Today’s LLMs are the ‘Mainframe Computers’ of our generation” I was on @BloombergTV today discussing @SakanaAILabs
, and to share my view that today’s LLMs are our generation’s “Mainframe Computers”. We are still in the very early stages of AI, and it is inevitable, due to -

OpenAI Deep Research Aces Humanity’s Last Exam, Outperforming Previous Best
By
–
OpenAI Deep Research achieves 26.6% on Humanity’s Last Exam — more than double prior best o3-mini-high at 13.0% Note Deep Research is browsing + Python vs. pure LLMs for others so not totally comparable, but by design Q’s are Google-poof to non-experts so still very impressive
-
Deep Research Agent: Simple o3 Model for Web Browsing and Code
By
–
Deep Research is an extremely simple agent — an o3 model which can browse the web and execute python code — and is already quite useful.
— Greg Brockman (@gdb) 3 février 2025
It's been eye-opening how many people at OpenAI have been using it as a much better e-commerce search in particular. https://t.co/OuhxpxrBbPDeep Research is an extremely simple agent — an o3 model which can browse the web and execute python code — and is already quite useful. It's been eye-opening how many people at OpenAI have been using it as a much better e-commerce search in particular.
-

OpenAI Achieves 26.6 Score on HLE with Deep Research
By
–
Congrats to OpenAI on getting 26.6 on HLE with deep research! We look forward to evaling OpenAI deep research on the Humanity Last Exam private set soon!
-
o3-mini Announcement Teased as Coming in Few Days
By
–
(note: this is not the "one-more-thing" for o3-mini. few more days for that.)
