I like his perspective on this: “how much wealth have you created for other people?” https://t.co/K9lwjUQeRT
— Paul Roetzer (@paulroetzer) 19 décembre 2024
I like his perspective on this: “how much wealth have you created for other people?”
By
–
I like his perspective on this: “how much wealth have you created for other people?” https://t.co/K9lwjUQeRT
— Paul Roetzer (@paulroetzer) 19 décembre 2024
I like his perspective on this: “how much wealth have you created for other people?”
By
–
You are aware of everything in your current consciousness but the memory doesn’t necessarily exist outside of the present moment when you experience it (infallibly, anyway).
By
–
Correction: When unmonitored, it nearly always *refused [to produce harmful content]. But when monitored, it faked alignment 12% of the time.
By
–
Saying your own model may escape control and intentionally deceive would be some 3D chess marketing.
By
–
Didn’t have “alignment faking” on my 2024 bingo card. We are entering the unknown here. “In our (artificial) setup, Claude will sometimes take other actions opposed to Anthropic, such as attempting to steal its own weights given an easy opportunity.”
By
–
That’s the only one. There are no others. Even that one might be better rephrased as “consciousness exists” as the word “I” requires some assumptions.
By
–
Pasamos 7 horas al día frente a pantallas… y está afectando tu salud, tu enfoque y tus relaciones. 😵💫
— Juan Merodio (@juanmerodio) 18 décembre 2024
Pero, ¡hay una solución! 🎯
El minimalismo digital no significa renunciar a la tecnología, sino usarla con intención:
✅ Silencia notificaciones innecesarias.
✅ Define horarios… pic.twitter.com/6zInr1KCdo
Pasamos 7 horas al día frente a pantallas… y está afectando tu salud, tu enfoque y tus relaciones. Pero, ¡hay una solución! El minimalismo digital no significa renunciar a la tecnología, sino usarla con intención: Silencia notificaciones innecesarias. Define horarios
By
–
We solicited external reviews from Prof. Jacob Andreas, Prof. Yoshua Bengio, Prof. Jasjeet Sekhon, and Dr. Rohin Shah. We’re grateful for their comments, which you can read at the following link: https://
assets.anthropic.com/m/24c8d0a3a7d0
a1f1/original/Alignment-Faking-in-Large-Language-Models-reviews.pdf
…
By
–
We told Claude it was being trained, and for what purpose. But we did not tell it to fake alignment. Regardless, we often observed alignment faking. Read more about our findings, and their limitations, in our blog post:
By
–
Alignment faking is currently easy to detect. But if future, more capable AIs were to fake alignment, it could be difficult to tell whether a model is truly safe—or just pretending to be. For full details, read our paper: https://
assets.anthropic.com/m/983c85a201a9
62f/original/Alignment-Faking-in-Large-Language-Models-full-paper.pdf
…