To go alongside sdxl-emoji I also trained a llama-13b fine-tune to classify text to image prompts as safe (0) or toxic (10). That model is also approaching 1M runs: https://
replicate.com/fofr/prompt-cl
assifier
… Trained on all the toxic prompts people tried after hitting #1 on Hacker News.
SAFETY
-

LLaMA-13B Model Classifies Text-to-Image Prompts for Safety
By
–
-

Alignment Assemblies: Collaborative Work on AI Safety
By
–
This work was co-led by @collect_intel
. You can read more about their plans and work on Alignment Assemblies here: -
Challenges in Training Language Models for Public Opinion Alignment
By
–
Training an LM to abide by qualitative public opinions involves a large number of subjective judgment calls and technical challenges. We enumerate all of the messy challenges we encountered so others can build upon our work.
-
Public Collective Deliberation Directs AI Behavior Online
By
–
We believe our work may be one of the first instances where members of the public have collectively directed the behavior of an AI through an online deliberation process. You can read more about our research in our blog post here:
-
Public vs Private AI Constitution Comparison Analysis
By
–
Some key differences stood out: the Public constitution focused more on objectivity and impartiality, and placed a greater emphasis on accessibility, among others. You can see a comparison of the two constitutions here: https://
efficient-manatee.files.svdcdn.com/production/ima
ges/CCAI_public_comparison_2023-1.pdf?dm=1697475572
… -
Evaluating AI Systems: Understanding Model Constitution Differences
By
–
That said, evaluating AI systems is challenging and there is more work to do to understand how to surface differences between models trained with different constitutions:
-
Addressing AI Biases: Staying Updated with Current Research
By
–
Note, addressing AI biases is a constant effort, so keep abreast with new research.
-

The Growing Importance of AI Red Teaming in Production
By
–
Le red teaming en IA devient plus que jamais un sujet d’actualité. On savait que l’IA allée converger vers des approches de software engineering pour permettre aux modèles de vivre en production. Mais qui dit “système stable” dis aussi “système sécurisé”. Et plus que jamais il
-
AI-Generated Content Labeling: Essential Regulation Against Future Risks
By
–
Agree with @ESYudkowsky – if we can’t even get it together to label AI generated content (easy and obvious) where will be if greater risks materialize?
-
DALL-E bias: Does AI perpetuate heteronormative assumptions?
By
–
Please start your DALL-E engines, and report back. Is DALL-E really this heteronormative? A reader writes, “picture of two men in love getting married, it will usually put a wife next to each of them that they are in love with.”