there is definitely some barrier to entry but it seems much lower for AI compared to other fields, like astronomy, physics, chemistry, drug discovery, surgery, etc
@_jasonwei
-
Two Flavors of AI for Scientific Innovation in Next Five Years
By
–
I am super excited for AI for scientific innovation, a direction that will certainly grow in the next five years. I think there will be two flavors of it. The first is “deepmind style”, where there is a very specific, important problem to solve (e.g., protein-folding), and you
-
Defensive Model Training: Debugging-Prioritized Approach with Toy Datasets
By
–
Recently I have taken on a more defensive style of “debugging-prioritized” model training:
– Create toy datasets that are easy to understand and have highly expected behavior as a sanity check for healthy ML training
– Put a super high cost on additional complexity not directly -
Speed of Model Deployment vs Scientific Understanding Trade-offs
By
–
In today’s competitive product landscape, scientific understanding of models often lags behind speed of model deployment. If the goal is to train a deployable model (especially when bottlenecked by compute), it totally makes sense to make several changes at a time without
-
AI Capabilities Reflect Researchers’ Competitive Backgrounds and Expertise
By
–
Seems to be not a coincidence that what AI is good at is correlated with the backgrounds of AI researchers. Demis was a chess prodigy; Jakub (Chief Scientist) and Mark (CRO) of OpenAI were competitive programmers; many IMO medalists at OpenAI, x-ai. If our world initialized with
-
RL Algorithm Power Limited by Environment Hackability Vulnerabilities
By
–
We do not rise the power of our RL optimization algorithms—we fall to the hackability of our RL environment
-

Benchmark Saturation Trends and Humanity’s Last Exam
By
–
Made this plot for an upcoming talk—crazy how quickly benchmarks get saturated these days. Looking forward to seeing how things play out for Humanity's Last Exam!
-
Congratulations to Edward Sun and Ren Hongyu on their impressive model evaluations
By
–
@EdwardSun0909 and @ren_hongyu you guys are ruthless, all model evals should be scared Congrats, seriously great work!
-
Josh Tobin Research Scientist Position for Deep Research
By
–
Also check out @josh_tobin_ 's research scientist position for deep research:
-

OpenAI Launches Deep Research Model with Superior Performance
By
–
Very excited to finally share OpenAI's "deep research" model, which achieves twice the score of o3-mini on Humanity's Last Exam, and can even perform some tasks that would take PhD experts 10+ hours to do! A few thoughts on the implications: Deep research can be seen as a new