And that’s before we flip into a version of training AI models through RL that has an infinite hunger for compute.
LLMS
-
Reasoning Models vs General LLMs Post-Training Techniques
By
–
Thanks for clarifying. E.g. in my view, "Pre-training, post-training, post-inference," would be too general because that's also the same flow for LLMs that are not specialized reasoning models. In a sense all reasoning techniques are post-training techniques because they all are
-

Groq Achieves Fastest Endpoint Speed at 1,566 Tokens Per Second
By
–
@GroqInc is offering the fastest endpoint at 1,566 output tokens per second! Thanks for benchmarking, @ArtificialAnlys
-

Chinese LLM Papers as Primary Learning Source for AI Research
By
–
“We are in this bizarre world where the best way to learn about LLMs is to read papers by Chinese companies. I do not think this is a good state of the world — US labs keeping their architectures and algorithms secret is ultimately hurting AI development in the US.” Thread
-

Interpretability Challenges in Large Language Models
By
–



Pour ceux qui veulent aller au-delà de la simple construction des LLMs et qui s’interrogent sur le problème de l’interprétabilité des résultats : Avec des dizaines (voire centaines) de milliards de paramètres, il est non seulement extrêmement difficile, mais aussi peu pertinent x.com/DFintelligence…
-

Intermediate model SFT data reuse for distillation
By
–
The way I read the paper, it was the intermediate model that generated the SFT data, and then they just reused that for distillation. But pls correct me if there some counterpoint to that in the paper.
-
Clarification on inference-time scaling versus test-time scaling terminology
By
–
Thanks for the comment! Actually, this part confused me a bit:
"whereas the methods you listed are ways we can induce test-time scaling – and so can't be meaningfully compared to "inference-time scaling"."
That's because "inference-time scaling" and "test-time scaling" should be -

10 Practical ChatGPT Task Ideas for Enhanced Daily Productivity
By
–
10 Best ChatGPT Tasks Ideas To Boost Productivity
Enhance your daily efficiency effortlessly
• Automate reminders
• Streamline workflows
• Manage schedules Click below to read more: https://
buff.ly/3WLG9I9 -

MultiChallenge: New Multi-Turn LLM Benchmark Released by Scale AI
By
–
Introducing MultiChallenge by @scale_AI – a new multi-turn conversation benchmark. Current frontier LLMs score under 50% accuracy (top: 44.93%). o1
Claude 3.5 Sonnet
Gemini 2.0 Pro Experimental Paper: http://
arxiv.org/abs/2501.17399
Leaderboard: http://
scale.com/leaderboard/mu
ltichallenge
… -

Andrej Karpathy Releases New Introductory LLM Video Series
By
–
A new video in Karpathy "Intro to LLMs for general audience" series dropped. This is a world gift. Thank you for all high quality videos, Andrej!!!
