crazy that in 2025 i can converse in 1000 tokens/sec on my single GPU machine with AI that’s world-class at math and programming but i still have to type. i can’t speak to it, at least not in low-enough latency to carry a conversation we don’t have this tech yet. why not?
@jxmnop
-
Current RL Paradigm Limitations Beyond Data Issues
By
–
ur right, data mixes sure aren't solved but the issues with the current RL paradigm go far beyond just the data mixes
-
Pretraining as Universal Compression: Learning to Simulate Nature
By
–
to pretrain is to learn to compress the universe. and as byproduct, learning to simulate all natural processes gather all the data you can. compress. backpropagate you don't see elegance in that?
-
AI Researchers Race to Perfect Reinforcement Learning Scaling
By
–
there’s a palpable tension in the air as hundreds of AI researchers (including me!) quietly work nights and weekends trying to figure out the “right way” to scale RL math & code are not the universe we will not rest until post-training is as clean and elegant as pre-training
-
USA AI breakthroughs: transformers, GPUs, diffusion models
By
–
happy birthday to the USA, the greatest country, and the origin of the following innovations: – Transformers
– Pre-training (web-scale next-token prediction)
– RLHF
– RLVR
– RL
– GPUs
– TPUs
– PyTorch
– word2vec
– reasoning models
– GANs
– diffusion models
– VLMs
– self-driving -
Waymo data reliability versus Tesla crowdsourced metrics
By
–
the waymo numbers are reliable the Tesla numbers are crowd sourced from users on a random site
-
Tesla vs Waymo: End-to-End vs Modular AI Agent Architecture
By
–
teslas drive an average of 400 miles before they require human intervention. waymos can drive for around 20,000 tesla vision is end-to-end. waymo is a modular system with many components, each an expert at one thing perhaps this tells us something about how to build AI agents
-
LLM Pretraining Evolution: From Crisis to Solved Problem
By
–
two or three years ago, this was a prevailing sentiment “training the largest LLMs is very hard. only a few people know how. everyone else is failing” it was hard to avoid huge loss spikes. now pretraining is a solved problem. what changed? do we just clean our data better?
-

Missed Naming Opportunity for New AI Model Paper
By
–
cool new paper/idea but a huge missed opportunity to name the model 5TPG