New blog post: How to Evaluate Jailbreak Methods: A Case Study with the StrongREJECT Benchmark
ETHICS
-
AI experts oppose SB 1047 bill call Newsom veto
By
–
Unfortunately, SB 1047 has passed the vote despite many AI experts like @drfeifei @ylecun and myself writing detailed articles on how fundamentally flawed it is. @GavinNewsom Please read our letter signed by @Caltech personnel, alumni, and others and veto this.
-
Social Apps Face Content Responsibility Test Online
By
–
This could be a real test of social apps’ long-held defence that the content that gets shared on their platforms isn’t their responsibility but users’.
-
AI layoffs strategy boosts stock performance
By
–
If you need to lay off staff, make sure to herald it as the benefits of AI. Very good for your stock.
-
AI in US Elections: Trump’s Strategy Targeting Taylor Swift Fans
By
–
La ia, como era de esperar, ha hecho su aparición en la campaña electoral de las elecciones presidenciales de los Estados Unidos.
— Juan Merodio (@juanmerodio) 28 août 2024
Qué te parece la estrategi de Donald Trump para poner a su favor a las seguidoras de Taylor Swift? pic.twitter.com/WqUKqt3lSjLa ia, como era de esperar, ha hecho su aparición en la campaña electoral de las elecciones presidenciales de los Estados Unidos. Qué te parece la estrategi de Donald Trump para poner a su favor a las seguidoras de Taylor Swift?
-

COAR: Scalable Predictive Component Attribution Method
By
–
In summary, COAR is a scalable method for estimating predictive component attributions that outperform prior approaches across models & tasks. COAR's attributions act as a counterfactual estimator, helping w/targeted model edits — from fixing errors to boosting subpopulation
-
COAR: Surgical Edits to Fix ML Model Errors
By
–
The approach, called COAR, allows us to make surgical "edits" to ML models. By intervening on the right components, we can fix errors, "forget" harmful labels, or boost performance on subpopulations that are underrepresented in data.
-
Safety Priority: Systems Check Before Launch Confirmation
By
–
We could have launched and it would have been fine, but safety is paramount, so better to check all systems again
-
AI Should Fact-Check User Content, Not Vice Versa
By
–
I want an AI that will fact-check what I write, not one whose writings I have to fact-check.