We covered a wide range of harm scenarios including: – Hacking and Malware
– Self harm and Suicide
– Harm to Children
– Illegal Activities
– Sexualized Content
– Graphic Violence … and more. See precise distribution below:
SAFETY
-

AI Harm Scenarios: Hacking, Safety, Violence, and Child Protection
By
–
-

Scale SEAL Leaderboard Evaluates AI Adversarial Robustness
By
–
1/ Scale is announcing our latest SEAL Leaderboard on Adversarial Robustness! Red team-generated prompts Focused on universal harm scenarios Transparent eval methods SEAL evals are private (not overfit), expert evals that refresh periodically http://
scale.com/leaderboard -

Global India AI Summit Addresses Safety, Access, Sustainability
By
–
During the recent #GlobalIndiaAISummit, the Innovation Workshop "Tackling AI Safety, Democratising AI and Sustainability" discussed key issues such as AI safety, democratizing AI access, and sustainability. The workshop emphasized the need for the international community to
-
Data Quality and Legal Optimization in AI Model Development
By
–
You mean that people will get used to the features? I still think there'll be a wave of legal "optimization" just like they are now optimizing performance. If you can build a model with 100% clean data, that has inherent value to those wanting to minimize risk (govt, bigcorp).
-
Model Collapse: Training AI on Synthetic Data Risks
By
–
9/ Model Collapse on Synthetic Data – investigates the effects of training models on recursively generated data; finds that training on model-generated content can cause irreversible defects where the original content distribution disappears; shows that the effect, referred to as
-
Existential concerns about technological future and societal collapse
By
–
Is there an option to just exit 2024 and go into cryogenic slumber until 2028 (or, should I say 2032 at this point – how bad is it gonna get before it gets good again?)
-
Location Data Privacy: Opt-In Requirements and Re-identification Risks
By
–
needs to be opt-in, not opt-out (if even)
selling location data anonymous makes it trivial to identify -
Bot or Expert: Moral and Intellectual Consistency in AI Systems
By
–
I think Ollie/Jenny is a bot, but if not, there is nothing but my admiration for your work and expertise in our interactions here! You have got things consistently right, both morally and intellectually. If there's any hint of frustration it's with the system itself…
-
OpenAI Launches ChatGPT Voice Mode After Safety Delays
By
–
In June, I reported that OpenAI delayed the launch of its new voice mode for ChatGPT to address safety issues; at the time, it said it would arrive in a month. Today @sama posted it'll drop next week (for a limited group of paid users). https://
bloomberg.com/news/articles/
2024-06-25/openai-delays-launch-of-voice-assistant-to-address-safety-issues?srnd=undefined
… -
Copyright Traps Paper Enhances Membership Inference Attacks Detection
By
–
Quoted in nice article by @Melissahei
, on a #ICML2024 paper on "copyright traps" by @matthieu_meeus @ffuuugor @ManuelFaysse @yvesalexandre Cool idea to enhance the efficacy of membership inference attacks to detect copyrighted data. Caution tho: similar to Glaze, not futureproof