I'm not 100% sure about that. As an example I was just browsing through the DCLM-baseline datamix (which is ~SOTA) and it is *terrible*. Compared to what I could in principle imagine. Major concessions are made in data quality to gather enough data quantity.
LLMS
-
The Quest for Ultimate High-Quality Pretraining Data for LLMs
By
–
Mildly obsessed with what the "highest grade" pretraining data stream looks like for LLM training, if 100% of the focus was on quality, putting aside any quantity considerations. Guessing something textbook-like content, in markdown? Or possibly samples from a really giant model?
-
Reasoning Models Revolution: New Scaling Laws Transform LLM Capabilities
By
–
"How many R’s are in strawberry?"
A year ago, ChatGPT would have guessed—or hallucinated. Today, it thinks. I just released a new video diving into the "revolution" of reasoning models and the new axis of scaling laws that's transforming how large language models operate. -
Favorite Open-Source LLM Community Poll
By
–
What’s your favorite open-source LLM these days? Drop your answer in the comments.
-

Neural Conversational Model Paper Turns 10 Years Old
By
–
"A Neural Conversational Model" is 10 years old, w/ @quocleix . TL;DR you can train a chatbot with a large neural network (~500M params!). Samples This paper was received with mixed reviews, but I'm glad all the critics are now riding the LLM wave https://
arxiv.org/abs/1506.05869 -
Using LLMs to Strengthen and Disperse Your Arguments
By
–
Playing with LLMs should be enough to disperse your argument for yourself
-
Embodied Web Agents: Bridging Physical-Digital Realms for Real-World Tasks
By
–
https://t.co/63LYHHBiEV pic.twitter.com/fNffHwnsN0
— AI Breakfast (@AiBreakfast) 20 juin 2025Meet Embodied Web Agents that bridge physical-digital realms. Imagine embodied agents that can search for online recipes, shop for ingredients and cook for you. Embodied web agents search internet information for implementing real-world embodied tasks. All data, codes and web
-
New Research Reveals Hidden Human Costs of ChatGPT Usage
By
–
But, now, read the new research on the hidden human cost of using ChatGPT:
-

Mistral Small 3.2 minor upgrade released
By
–
Mistral AI released Mistral Small 3.2 with a minor upgrade for instruction following, repetition errors and function calling!
-

Mistral Small 3.2 Improves Instruction Following and Reduces Repetition
By
–
Introducing Mistral Small 3.2, a small update to Mistral Small 3.1 to improve: – Instruction following: Small 3.2 is better at following precise instructions
– Repetition errors: Small 3.2 produces less infinite generations or repetitive answers – Function calling: Small