My biggest takeaways from @benedictevans
: 1. We’re in 1997 for AI—it’s as big a deal as the internet or mobile, and only as big a deal as the internet or mobile. We’re at the stage where most stuff kind of doesn’t work yet, most of what people will build hasn’t been built, and
RESEARCH
-
AI in 1997: As Big as Internet or Mobile, Still Early
By
–
-

Runway announces London as European HQ and $100M AI investment
By
–
Today we're announcing London as Runway's new European headquarters and our newest research hub focused on general world models. Over the next 18 months, we plan to invest $100M into the UK AI ecosystem, and that figure will more than double through 2028 as we scale our European
-

Paper addresses call for better benchmarks in private ML
By
–
I'm excited since this paper addresses a call-to-arms in our #ICML2024 Best Paper with @florian_tramer and Nicholas Carlini, where we advocated for better benchmarks in private ML: https://
x.com/thegautamkamat
h/status/1603383883126669312
… 7/n -

DP synthetic data benchmarks: hard even with large privacy budgets
By
–
These tasks are hard/impossible to zero-shot, rather easy without privacy, but surprisingly hard even with large privacy budgets (ε = 100)! This room to grow means we can really measure progress made by new DP synthetic data benchmarks. 6/n
-

New ContinuousBench tasks replace saturated DP-synth benchmarks
By
–
3. Current DP-synth methods shouldn't perform too well: else, there's no room to distinguish new and better techniques. Classic benchmarks used for DP synth (e.g., IMDb, OpenReview) are effectively saturated. Our new ContinuousBench tasks (Geminon and News) satisfy 1-3. 3/n
-
ContinuousBench prevents benchmark leakage and measures learning from data
By
–
Benchmarks may leak. They might also measure something besides what we want to measure: how well does a model learn from data? ContinuousBench will be released periodically, preventing leakage, and tasks are chosen so that information is exclusively in the data. 4/n
-
Benchmarking DP synthetic data: zero-shot and real data training
By
–
So you want to see if your DP synthetic data method is actually any good. What makes a good benchmark? 1. Zero-shot performance should be low: the method should measure learning from the actual data;
2. Training on real data should work: learning should be possible; and… 2/n -

ContinuousBench: Hard Leakage-Proof DP Synthetic Text Benchmark
By
–
Does DP synth text transfer useful knowledge or just superficial style mimicking? Existing benchmarks: saturated Introducing ContinuousBench: a hard (curr methods fail at ε=100! ) & leakage-proof benchmark for DP synth text! Followup to our #ICML2024 best paper 1/n
-
NVIDIA Launches Cosmos Coalition for Open World Models
By
–
7/ NVIDIA also launched the Cosmos Coalition – with Black Forest Labs, Runway, Skild AI, Agile Robots, LTX and Generalist – to push open world models forward together. The bet: world models are becoming the intelligence layer for robots and AVs, and NVIDIA wants the open
-
Cosmos 3 tops open leaderboards for text, video, physics, and robotics
By
–
5/ And it's not a demo. Cosmos 3 tops the open leaderboards: #1 open model on Artificial Analysis for text→image AND image→video
#1 on Physics-IQ for physics accuracy — ahead of Sora 2
Leads PAI-Bench overall, ahead of Veo 3.1
#1 robot policy on RoboArena
Open weights beating