I'm excited since this paper addresses a call-to-arms in our #ICML2024 Best Paper with @florian_tramer and Nicholas Carlini, where we advocated for better benchmarks in private ML: https://
x.com/thegautamkamat
h/status/1603383883126669312
… 7/n
@thegautamkamath
-

Paper addresses call for better benchmarks in private ML
By
–
-

DP synthetic data benchmarks: hard even with large privacy budgets
By
–
These tasks are hard/impossible to zero-shot, rather easy without privacy, but surprisingly hard even with large privacy budgets (ε = 100)! This room to grow means we can really measure progress made by new DP synthetic data benchmarks. 6/n
-
ContinuousBench prevents benchmark leakage and measures learning from data
By
–
Benchmarks may leak. They might also measure something besides what we want to measure: how well does a model learn from data? ContinuousBench will be released periodically, preventing leakage, and tasks are chosen so that information is exclusively in the data. 4/n
-

New ContinuousBench tasks replace saturated DP-synth benchmarks
By
–
3. Current DP-synth methods shouldn't perform too well: else, there's no room to distinguish new and better techniques. Classic benchmarks used for DP synth (e.g., IMDb, OpenReview) are effectively saturated. Our new ContinuousBench tasks (Geminon and News) satisfy 1-3. 3/n
-
Benchmarking DP synthetic data: zero-shot and real data training
By
–
So you want to see if your DP synthetic data method is actually any good. What makes a good benchmark? 1. Zero-shot performance should be low: the method should measure learning from the actual data;
2. Training on real data should work: learning should be possible; and… 2/n -

ContinuousBench: Hard Leakage-Proof DP Synthetic Text Benchmark
By
–
Does DP synth text transfer useful knowledge or just superficial style mimicking? Existing benchmarks: saturated Introducing ContinuousBench: a hard (curr methods fail at ε=100! ) & leakage-proof benchmark for DP synth text! Followup to our #ICML2024 best paper 1/n
-

Don’t let AI steal your thinking and personality in talks
By
–
In the last 48h:
– Jr researcher asked me wheter to use AI in making talks
– Saw two talks, with AI {slop, enhanced} slides Collected my thoughts and wrote a post. Tl;dr: don't steal your own thinking, don't remove *you* from your talks. Also, give a &#@% about your talks. -
New generative model research uses low-rank Nyström approximation for faster training
By
–
Ali did some amazing work on the hottest new generative model: drifting models (from Mingyang Deng @Goodeat258 et al., out of Kaiming He's group).
— Gautam Kamath (@thegautamkamath) 14 mai 2026
Speeds up training a lot using low rank Nyström approximation. Check out Ali's full thread. Paper and code available! https://t.co/MDSo47g4rUAli did some amazing work on the hottest new generative model: drifting models (from Mingyang Deng @Goodeat258 et al., out of Kaiming He's group). Speeds up training a lot using low rank Nyström approximation. Check out Ali's full thread. Paper and code available!
-
Accepted papers for COLT 2026 announced
By
–
Accepted papers for #COLT2026: https://
learningtheory.org/colt2026/accep
ted.html
… -
Research Incentives: Quality Over Quantity in Academia
By
–
Is that unfortunate? I feel like it sets the incentives as they should be: to try to make deep contributions rather than spamming papers.