It’s is a burgeoning area of interest on how yo make your application full proof to prompt leaks or any kind of prompt injection attacks. Some remedies are making sure that people can’t easily get to your source prompt (secret sauce) by tweaking with the input.
SAFETY
-
Simulating Trump vs ordinary people: AI capability differences
By
–
Like he says ppl like Trump would be easy to simulate while I say normal people are easy to simulate
-
Bypassing AI Detection: Natural Text and Watermarking Constraints
By
–
I'm not sure about that. E.g. tricking the ZeroGPT algo means actually writing more natural text (so it'll be better, not worse), and tricking the watermarking means not being limited by the constraints of the watermarking algo.
-
Invisible Watermarking Systems Easily Defeated by Adversarial Methods
By
–
Por ejemplo, me estáis compartiendo un vídeo donde se propone marcar los textos de ChatGPT de forma invisible para poder detectar que es artificial. El problema, como comenta aquí Jeremy, es que estos sistemas una vez creados, son muy fáciles de superar.
-

Blurring Lines Between AI and Human-Generated Content Online
By
–
Por otro lado, las fronteras que intentamos poner entre contenido AI/Humans son cada vez más difusas, creando situaciones delicadas como estas, donde en Reddit a un usuario se le ha baneado bajo la sospecha de que su arte (que él dice es original suyo) puede estar hecho con IA…
-
Can AI-Generated Content Be Reliably Detected?
By
–
¿Cómo podemos detectar si una imagen o un texto ha sido creado por una IA? ¿Se puede usar otra IA? Esto es algo que muchos me preguntáis y malas noticias: NO SE PUEDE Al menos por ahora no hay método satisfactorio que te permita identificar si algo es artificial… [1/n]
-
AI Struggles with Accurate Character and Number Counting
By
–
Yeah, it's just not great at counting. I think it sort of knows approximately how long a number is, but you'll always have to police the character counts. Interesting especially on the tweet length. That makes total sense.
-

Understanding Deep Learning Overfitting Through Mechanistic Analysis
By
–
We have little mechanistic understanding of how deep learning models overfit to their training data, despite it being a central problem. Here we extend our previous work on toy models to shed light on how models generalize beyond their training data. https://
transformer-circuits.pub/2023/toy-doubl
e-descent/index.html
… -
Watermarking Limitations in Democratized LLM Training Era
By
–
In the last sentence of his reply he says "in a world where" people other than OpenAI can train an LM, this watermarking fails. I guess he doesn't realise we're in that world right now, when it comes to restating/styling/summarising prose.