AlphaOne: Reasoning Models Thinking Slow and Fast at Test Time Zhang et al.: https://
arxiv.org/abs/2505.24863 #ArtificialIntelligence #DeepLearning #MachineLearning
LLMS
-

AlphaOne: Reasoning Models with Slow and Fast Thinking
By
–
-
AI Understanding Physical World for Scientific Discovery
By
–
It was an honor to be part of @Google IO Dialogues stage with James Manyika, Pushmeet Kohli and Joëlle Barral and talk about AI+Science. I talked about how AI needs to understand the physical world in order to make new scientific discoveries. While LLMs can come up with new
-
Method-Driven Research and AI: The Hammer Paradigm Shift
By
–
There are traditionally two types of research: problem-driven research and method-driven research. As we’ve seen with large language models and now AlphaEvolve, it should be very clear now that total method-driven research is a huge opportunity. Problem-driven research is nice because you have a consistent and specific goal. The goal is usually virtuous, so it feels good to have a mission and identity. However, it just doesn’t work due to The Bitter Lesson. Basically everything in classical NLP (machine translation, summarization, chatbots) lost to simple scaling. ChatGPT is a prime example—it used nothing from chatbot research and certainly wasn’t the intended end goal of OpenAI’s 2022 research program, but was a huge hit because someone (John Schulman et al) figured out the right way to package large language models as a product. Method-driven research feels less stable because you’re constantly searching for problems and you have to be opportunistic. But I believe AI will allow method-driven research to dominate progress in most fields of science, one-by-one. The latest method (or “hammer”), as we’ve seen in AlphaEvolve, is ruthless search and optimization against a reward function (whether this requires RL or not is a separate discussion). Things that problem-driven researchers have been trying to solve for a long time like the kissing number problem will become nails hit by the hammer. Eventually the hammer will become bigger, stronger, and more general and will hit more and more nails. So a very important meta-skill for the next decade will be knowing how to create the right environments to use The Hammer. Ironically, the problem-driven researchers, who by definition are experts in a specific problem, are well-positioned to create these environments. If, that is, they can put down their egos and pick up the hammer.
→ View original post on X — @_jasonwei, 2025-06-02 19:30 UTC
-
LLM Limitations: Human Scaffolding Requirements for Output Convergence
By
–
I think this distinction sometimes elides people who are quite steeped in day to day LLM usage, because you can imagine a system that you can cobble together which would converge on correct output after you have spent hundreds of hours on being human scaffolding for it.
-

Alex Albert Judges Hackathon, Anticipates Impressive Claude 4 Projects
By
–
Very excited to be one of the judges for this hackathon! Expecting to see some pretty amazing Claude 4 projects
-
LLMs Image Generation Enables Quick Concept Feasibility Communication
By
–
(I think LLMs doing image generation is a bit of an unlock here because it can more quickly communicate “OK so is something in this direction possible enough for you to draw sketches or am I asking for a functioning fusion reactor with inlaid Japanese lacquerwork.”)
-
Choosing Between o3 and 4o Based on Task Importance
By
–
Got it! I think I make the decision of whether something is important (and I'm willing to wait) or not that important (and I just want to get a fast sense) and that basically determines if I go to o3 or 4o. It's conceptually easy to just make a binary decision. I'll try it more!
-
Karpathy Endorses Perplexity for Search and Quick Summaries
By
–
I really like Perplexity and use it for anything "search-like", though other LLM providers now include search. It's fast and works great, and is also very useful for quick summaries of whatever trending topics there are. (I'm an investor fyi, but <3 for reals).
-

O3 Reasoning Model Superiority for Important Tasks
By
–
An attempt to explain (current) ChatGPT versions. I still run into many, many people who don't know that:
– o3 is the obvious best thing for important/hard things. It is a reasoning model that is much stronger than 4o and if you are using ChatGPT professionally and not using o3 -
Windows AI PCs: 40+ TOPS Benchmark, Local Models, and New Features
By
–
We also covered: – Why Windows established a 40+ TOPS benchmark for AI PCs
– How small language models like Phi-4 are unlocking local reasoning
– Why Recall, Cocreator & other new features are just the beginning Read the full Q&A here: