move over meta, the true biggest benefactor of open source machine learning is CHANEL
@jxmnop
-
Meta’s 15T Token Training Strategy Preference
By
–
would prefer they just trained on 15T tokens like meta
-

Weights Biases Data Visualization Limitations Critique
By
–
weights & biases is great software but has anyone ever learned anything from one of these graphs? can't even see a correlation between two variables if their columns aren't next to each other
-
Phi Training Approach Criticism: Quality Data vs Quantity Debate
By
–
in my opinion, the Phi approach to training language models is just wrong • i'm not convinced that training on less (albeit "higher-quality") data is better than training on as much data as possible
• i'm not convinced that training on synthetic data ever works better than -
Common Research Pitfall: Confounding Variables in Experiments
By
–
common dark thought pattern in research
> run baseline experiment
> change thing A
> also change thing B
> run new experiment
> collect results
> "wow, thing A works!" -

Exponential Growth Perception During Sigmoid Curve Inflection
By
–
yearly reminder everything looks exponential from the middle of a sigmoid
-

Exponential Growth Perception in Sigmoid Curves
By
–
yearly reminder everything looks exponential from the middle of a sigmoid
-
MMLU Performance Shows Exponential Growth Over Time
By
–
I had the same thought when I was listening to this — I think it’s exponential if u plot MMLU (model performance) vs time (year achieved)
-
Compute Optimal Inference Setups Explained
By
–
well, it's compute optimal for some inference setups, I think that's the idea; it's not optimal for training though
-

Scaling Laws Evolution: 8B Model Trained on Fifteen Trillion Tokens
By
–
you're telling me an 8B param model was trained on fifteen trillion tokens? i didn't even know there was that much text in the world really interesting to see how scaling laws have changed best practices; GPT-3 was 175 billion params and trained on a paltry 300 billion tokens
