AI Dynamics

Global AI News Aggregator

About

Why models trained on ChatGPT don’t all excel

If DeepSeek-V3 is good because it trained on ChatGPT (which of course it did), why isn’t Grok amazing? Why isn’t *every* model amazing? Why spend 95% of compute pre-training a new model (which equals 405B on Pile-test btw) if the secret sauce is ~fOrBiDdEn~DaTa~ in the last 5%?

→ View original post on X — @goodside