Leaks look credible, Llama 3 405B is GPT-4o class It even outperforms GPT-4o (72.55%) and Claude 3.5 Sonnet (72.83%) in terms of MMLU PRO. The new 70B also looks insane with a significant boost of performance compared to the previous version. Note that these evals are not
LLMS
-

Machines Understanding Natural Language in Intelligent Factories
By
–
What if machines could understand and communicate with us effortlessly in our own language? It was inspiring to host @AmolAdgaonkar at @Microsoft and explore the era of Intelligent Factories during CXO Spice. Amol highlighted a pivotal shift: machines now converse in natural
-
Explaining Attention Mechanism to Non-Technical Audience
By
–
I just explained the essence of “attention is all you need” to my dad
-
Large Models Training Smaller Models: Data Generation Strategy
By
–
I think it only works when you have a larger model produce data for a smaller one. Otherwise the large model -> same large model, without any other external signals, will likely just learn something very similar
-

KL Divergence: The Foundation for Modern Language Models
By
–
http://
blog.alexalemi.com KL is All You Need https://
bit.ly/3Ll57rQ
#AI #MachineLearning #DeepLearning #LLMs #DataScience -

New favorite LLM test reveals inconsistent performance across SOTA models
By
–
Wow, this has just become my favorite LLM test. I missed that this doesn't work but it really doesn't, even for SOTA LLMs. Seems to be a bit hit and miss, e.g. with GPT4o which failed 1/3 times, Claude failed 3/3 times.
-
Llama-3.1-405B for Training Dataset Generation
By
–
Who's going to use Llama-3.1-405B to generate the best training dataset for small models?
-

Small Models Large Context Windows: The Winning LLM Formula
By
–
The winning formula for LLMs is going to be (a) small model, (b) large context window, (c) strong reasoning ability. The quality cost trade off makes this clear: https://
x.com/swyx/status/18
15037679014388172
… Large context window will enable a lot of applications (prompt instructions, RAG) as long -
Open Source to Dominate: Smaller Models Over Large Language Models
By
–
This also means that Open Source is going to dominate in the future – training and serving smaller models is more accessible to many. The large models are going to useful for generating and filtering data, while the smaller models (e.g., 4o-mini) are the ones we want deployed.
-

GPT-2 Sized Model Demonstrates Universal Physical Understanding
By
–
Building AI with universal physical understanding will have a tremendous impact. We have taken the first steps towards this by building a ~GPT-2 sized model that shows that a broad model can capture a wide range of physical phenomena. https://
arxiv.org/pdf/2403.03542 We have scaled the
