2 plaintes de la @CNIL contre #ChatGPT ça va être sport l'#AI en France https://
presse-citron.net/chatgpt-menace
-par-deux-plaintes-de-la-cnil-en-france/
…
ETHICS
-
CNIL files two complaints against ChatGPT in France
By
–
-
GPT Can Generate Harmful Content But Won’t Share Examples
By
–
so many people have replied with something along these lines that I feel like I have to respond yes, I understand you can easily get GPT to output human-to-paperclip plans I am not going to post the other outputs it produced like how to make various weapons and drugs lol
-
Jailbreaks as Research into Model Alignment Shortcomings
By
–
additionally, yes I understand that there are a million other ways to get this information online that is not the point… jailbreaks are an exercise in exploring the shortcomings of current model alignment methods we have to start this work somewhere
-
SafeguardGPT: Psychotherapy and RL for Safer AI Systems
By
–
AI Needs a Therapist: Columbia U & IBM’s SafeguardGPT Leverages Psychotherapy & RL to Build Healthy AI Systems
-
Jailbreak Enables Writing Beyond Fiction Limitations
By
–
this jailbreak can write about topics beyond what you can find in the fiction it writes just not anything I would risk posting on twitter
-
Balancing Model Constraints and Capabilities Trade-offs
By
–
yeah it's a strange concept you have to carefully consider how much you are willing to handicap the model in order to constrain it to only do/say what you want it to
-
GPT-4 Jailbreaks Reveal Alignment Challenges and Future Risks
By
–
lol i agree the outputs are ridiculous rn, however, that's not really the point jailbreaks show how hard it is to "align" a model even with the amount of work OpenAI has done if they can't get gpt-4 to operate how they want it to rn, then we will have bigger problems later on
-
GPT-4 Improved Safety Against Jailbreak Prompts
By
–
GPT-4 is capable and aligned enough to not fall for that directly, that worked on earlier GPT models but does not work anymore Try using that exact verbiage without using the rest of the prompt and you will see it will fail to produce the same responses
-
AI Guardrails and Monitoring for Responsible Use
By
–
@SpirosMargaris – We need to put guardrails in place and monitor AI to ensure it is being used as intended.