Top AI stories today: – Nvidia threads agents across the stack
– Bernie Sanders seeks a public AI stake with new bill
– Turn Claude sessions into skills with a daily audit
– Hackers access IG accounts by…asking Meta AI?
– 4 new AI tools, community workflows, and more
SECURITY
-

Top AI stories today and updates
By
–
-
AI replaces Trust and Safety before AI-driven theft wave
By
–
Replacing Trust and Safety with AI right before an AI-driven account theft wave is the kind of timing you couldn't script.
-
Merge launches Agent Handler for Employees, secure AI integration with company software.
By
–
🆕: Merge just launched Agent Handler for Employees — the safest way to connect AI to your company software.
— Futurepedia – Learn to Leverage AI (@futurepedia_io) 2 juin 2026
✅ Connects Claude, ChatGPT, Cursor, Copilot & more to any third-party system
✅ IT imports employees → maps them to approved tools
✅ Built-in data loss prevention… pic.twitter.com/RgQvgShtNJ: Merge just launched Agent Handler for Employees — the safest way to connect AI to your company software. Connects Claude, ChatGPT, Cursor, Copilot & more to any third-party system IT imports employees → maps them to approved tools Built-in data loss prevention
-
Interrupt keynote on sandboxes for safe agent code execution
By
–
.@MukilLoganathan’s Interrupt keynote on Sandboxes. https://t.co/oddQOs0Q6O
— LangChain (@LangChain) 1 juin 2026
In 20 minutes, you’ll learn how to run agent code safely.
Isolated from your runtime, with network controls, persistent state, and snapshot/restore when things go wrong. pic.twitter.com/g2Pvzi824D.
@MukilLoganathan
’s Interrupt keynote on Sandboxes. https://
youtu.be/IIchUA5T3gs In 20 minutes, you’ll learn how to run agent code safely. Isolated from your runtime, with network controls, persistent state, and snapshot/restore when things go wrong. -

Paper addresses call for better benchmarks in private ML
By
–
I'm excited since this paper addresses a call-to-arms in our #ICML2024 Best Paper with @florian_tramer and Nicholas Carlini, where we advocated for better benchmarks in private ML: https://
x.com/thegautamkamat
h/status/1603383883126669312
… 7/n -

DP synthetic data benchmarks: hard even with large privacy budgets
By
–
These tasks are hard/impossible to zero-shot, rather easy without privacy, but surprisingly hard even with large privacy budgets (ε = 100)! This room to grow means we can really measure progress made by new DP synthetic data benchmarks. 6/n
-

New ContinuousBench tasks replace saturated DP-synth benchmarks
By
–
3. Current DP-synth methods shouldn't perform too well: else, there's no room to distinguish new and better techniques. Classic benchmarks used for DP synth (e.g., IMDb, OpenReview) are effectively saturated. Our new ContinuousBench tasks (Geminon and News) satisfy 1-3. 3/n
-
Benchmarking DP synthetic data: zero-shot and real data training
By
–
So you want to see if your DP synthetic data method is actually any good. What makes a good benchmark? 1. Zero-shot performance should be low: the method should measure learning from the actual data;
2. Training on real data should work: learning should be possible; and… 2/n -

ContinuousBench: Hard Leakage-Proof DP Synthetic Text Benchmark
By
–
Does DP synth text transfer useful knowledge or just superficial style mimicking? Existing benchmarks: saturated Introducing ContinuousBench: a hard (curr methods fail at ε=100! ) & leakage-proof benchmark for DP synth text! Followup to our #ICML2024 best paper 1/n
-
Mythos cybersecurity capabilities, not benchmarks, put it on White House watch list
By
–
Remember that Mythos isn't just a chatbot. The cybersecurity capabilities are what put it on the White House's watch list, not benchmark scores.