7/ We see strengths (Claude 3.5 w/coding, Gemini 1.5 w/Vision, Multilingual) from the models that are likely driven by post-training and data strategies. This is contrary to the prevailing beliefs 1 yr ago, where pre-training was believed to be the core competitive driver.
LLMS
-
Post-Training Data Strategy: Key Competition Arena for LLMs
By
–
6/ Post-training data strategy is becoming a key arena for competition. Llama3.1 paper had 15 pg about post-training data (vs. 12 pg on pre-training), driving these capabilities: – Code
– Multilinguality
– Math/Reasoning
– Long-Context
– Tool Use
– Factuality
– Steerability -
Post-Training Data Strategies: SFT, RLHF, and DPO Approaches
By
–
4/We are also seeing remarkably similar data strategies for post-training from most labs at this point (at least what was published from Meta+Apple): – Hybrid data SFT, RLHF, & DPO setups
– Synthetic data on code and math
– Post-training data for most important capabilities -
H100 GPU Rollout Explains Timing of Recent AI Model Releases
By
–
3/The reason these are all so close together timing-wise is that every lab got their H100s at roughly the same time. They each struggled with early issues with the H100s last fall, and the big H100 clusters all started training this spring. Voila, 5-6 months later, big models!
-
Seven Major AI Models Released in Three Months
By
–
2/We've seen 7 major models from top labs in the last 3mo: May:
– GPT 4o
– Gemini 1.5 Pro June:
– Claude 3.5 Sonnet July:
– Llama 3.1
– Mistral Large 2
– GPT-4o Mini August:
– Gemini 1.5 0801 Each of these models has been incredibly competitive—each world-class in some way. -
Gemini 1.5 Pro Tops LMSYS: Google’s Compute Edge Advantage
By
–
1/Gemini 1.5 Pro 0801 is the new best model (tops LMSYS, SEAL evals incoming) Key considerations
1—OpenAI, Google, Anthropic, & Meta all right ON the frontier
2—Google has a long-term compute edge w/TPUs
3—Data & post-training becoming key competitive drivers in performance -

Gemini 1.5 Pro Tops AI Leaderboard in Latest Developer Preview
By
–
Never seen a competitive leaderboard that I didn't like Congrats to the Gemini team on ranking no.1 with our latest improved Gemini 1.5 Pro developer preview model, which you can try on AI studio now!
-

New KV Cache Framework Integrates Latest Research Papers
By
–
For those of you that can't keep up with all the new stuff from @answerdotai
, I've got bad news–we have more! @GriffinAdams92
's first project is a KV cache framework that brings the latest research papers together. (And rather than write a blog post, he wrote an epic tome.) -

OpenDevin: Open Source Generalist LLM Agents Framework
By
–
Open source generalist LLM Agents! Excited to have author @xingyaow_ on alphaXiv this week to answer questions about his recent paper OpenDevin.
-

Google DeepMind’s Gemini 1.5 Pro Tops LLM Leaderboard
By
–
@GoogleDeepMind takes the lead with Gemini 1.5 Pro for the first time on the @lmsysorg leaderboard! We're excited to try this model out!