2/We've seen 7 major models from top labs in the last 3mo: May:
– GPT 4o
– Gemini 1.5 Pro June:
– Claude 3.5 Sonnet July:
– Llama 3.1
– Mistral Large 2
– GPT-4o Mini August:
– Gemini 1.5 0801 Each of these models has been incredibly competitive—each world-class in some way.
@alexandr_wang
-
Seven Major AI Models Released in Three Months
By
–
-
Gemini 1.5 Pro Tops LMSYS: Google’s Compute Edge Advantage
By
–
1/Gemini 1.5 Pro 0801 is the new best model (tops LMSYS, SEAL evals incoming) Key considerations
1—OpenAI, Google, Anthropic, & Meta all right ON the frontier
2—Google has a long-term compute edge w/TPUs
3—Data & post-training becoming key competitive drivers in performance -
Tech Executive Cancels TechCrunch Talk Over Critical Coverage
By
–
Given a TechCrunch writer’s recent ad hominem attacks while covering our MEI policy, I’ve decided to call my TechCrunch Disrupt session off. Yes, “all the way off.”
-
Gemini Team Achievements Recognized by Google DeepMind
By
–
Great work to the Gemini team and @GoogleDeepMind
-

Gemini 1.5 Pro tops AI safety harm classification leaderboards
By
–
We classified harms into High Harm and Low Harm. Congratulations to @GoogleDeepMind Gemini 1.5 Pro (post-IO) for being top of the list on both leaderboards!
-

AI Harm Scenarios: Hacking, Safety, Violence, and Child Protection
By
–
We covered a wide range of harm scenarios including: – Hacking and Malware
– Self harm and Suicide
– Harm to Children
– Illegal Activities
– Sexualized Content
– Graphic Violence … and more. See precise distribution below: -

Scale SEAL Leaderboard Evaluates AI Adversarial Robustness
By
–
1/ Scale is announcing our latest SEAL Leaderboard on Adversarial Robustness! Red team-generated prompts Focused on universal harm scenarios Transparent eval methods SEAL evals are private (not overfit), expert evals that refresh periodically http://
scale.com/leaderboard -
LLM Data Wall Challenge and Synthetic Data Solutions
By
–
Definitely a real issue with LLM development. The "data wall" + how challenging it is to generate new data is a major elephant in the room. That being said, not insurmountable:
– Hybrid data where human experts produce dramatically more data utilizing synthetic methods is going -
Synthetic Data Debt: Model Quality Degradation Risk
By
–
7/ watch this space closely. My prediction is that model developers who are not careful about training on synthetic data with no information gain will find their models getting steadily stranger and dumber over time. Synthetic data accumulates a debt with the model that must be
-
Hybrid Data Over Pure Synthetic Data for AI Models
By
–
6/ This is why I personally believe in HYBRID DATA over pure synthetic data. Synthetic data must be generated utilizing some source of new information for the model, either: (1) using a seed of real world data
(2) human experts
(3) formal logic engine Hybrid data is the real