5/ While the experiments done here were on small scale models (100M parameters), the fundamental effects we’re seeing here are likely to show up over time on the larger models too. For example, most of today’s models can’t produce a blog post in the style of Slate Star Codex,
@alexandr_wang
-

Three Sources of Error in Synthetic Data Training Generations
By
–
4/ There are three sources of error that accumulate from successive generations of synthetic training 1) statistical approximation error
2) functional expressivity error
3) functional approximation error Roughly speaking, each time you train the model on data generated from the -
Synthetic Data Training Risks Model Collapse Long Term
By
–
3/ This core idea is very important to pay attention to: Synthetic data can create a short-term boost in eval results, but you will pay for it later with model collapse! You accumulate debt with mangling the model that starts invisible, and is very hard to repay.
-

Synthetic Data Training Limitations and Mode Collapse in Self-Distillation
By
–
Training on pure synthetic data has no information gain, thus there is little reason the model *should* improve. Oftentimes when evals go up from “self-distillation”, that might be from some more invisible tradeoff, i.e. mode collapse in exchange for individual eval improvement
-

Model Collapse in AI: Recursive Synthetic Data Training Risks
By
–
1/ New paper in Nature shows model collapse as successive model generations models are recursively trained on synthetic data. This is an important result. While many researchers today view synthetic data as AI philosopher’s stone, there is no free lunch. Read more
-
GPT-4 Turbo Outperforms GPT-4o on Key Evaluation Dimensions
By
–
indeed. we haven’t finished our 4o mini eval, but our evals indicate that 4 turbo is still better than 4o on many important dimensions
-

Scale AI Partnering with Meta for Llama 3.1 Enterprise Adoption
By
–
Scale AI will also be partnering very closely with @Meta on Enterprise adoption of Llama 3.1. Zuck mentioned Scale AI in his memo on Llama3.1 open-sourcing as one of the key partners for enterprise adoption and custom Llama models.
-
Scale Data Foundry Advances Llama-3.1 Performance with Frontier Data
By
–
Lastly, as I've spoken about before, frontier data generation is critical to driving frontier model performance. Scale Data Foundry was utilized to generate frontier data (SFT & RLHF data) to push the performance of Llama-3.1 We are excited to continue pushing SOTA with Meta!
-

Meta Releases Llama 3.1 405B with Scale AI Partnership
By
–
1/Meta just released Llama3.1 405B! @scale_AI partnered deeply with @Meta on this release: SEAL Evaluations: Based on our evals on IF on Math #4 on Coding Enterprise partnership for custom Llama models Data Foundry partnership on RLHF & SFT
-
Llama 3.1 405B Achieves Top Performance on SEAL Leaderboard
By
–
We evaluated Llama3.1 405B Instruct on SEAL Leaderboard. As a reminder, our evals are:
PRIVATE (no overfitting)
EXPERT EVALUATED (trustworthy)
EVOLVE (no saturation) Our results show Llama3.1 is top notch:
Instruction Following
Math
#4 Coding