A great talk on "learning to reason with LLMs" by Naom Brown(
@polynoamial
). – Scale compute at inference time, not just train time. Let the model "think longer" using RL to improve reasoning. More like System 2 thinking, but for LLMs. – Optimize the "chain of thought" with
LLM Reasoning: Scaling Inference Compute and Chain of Thought Optimization
By
–
