LLaMA 3 70B cleanly beats Bixtral 8x22B.
@mattshumer_
-

LLaMA 3 400B Matches Claude 3 Opus Performance
By
–
The craziest LLaMA 3 reveal: The 400B+ version of the model is **on par with Claude 3 Opus**, and it's still training. Soon, we'll have a better-than-Opus, fully open-source model. The implications are huge.
-

LLaMA 3 70B Outperforms Claude 3 Sonnet Cost-Effectively
By
–
Holy shit. LLaMA 3 70B cleanly beats Claude 3 Sonnet. Small enough to host at scale without breaking the bank.
-
Training LLMs: Avoiding Default Mode Through Data Curation
By
–
First you’d teach the model quite a bit from recent content (w/ tons of LLM outputs in it it’s likely, if we were to train this in normally, the LLM will likely get stuck in that default LLM ‘mode’ we know so well + make it harder to break out of this w/ post-training) So
-
Training Strategy: Pre-2021 Data Priority Over AI-Generated Content
By
–
With all the AI-generated content flooding the web There might be something to first training on content from 2021-on And then continuing to train on pre-2021 content
-
Llama 3 Open Source Release Announced Within 24 Hours
By
–
less than 24 hours till llama 3 is open-sourced
-
Anthropic Reports Claude 3 Quality Drift Update
By
–
Update from the Anthropic team re: Claude 3 quality drift:
-
Fine-tuning Strategy: Sequential Layer Training Approach
By
–
Has anyone tried fine-tuning half of the layers in one run, and then the other half in another? cc @teknium @winglian @erhartford
