Worth noting that just because the system prompt says "based on the GPT-4 architecture" doesn't mean that the model is actually based on GPT-4! The goal of a system prompt is to influence the model to behave in certain ways, not to give it truthful information about itself
SAFETY
-
Corporate Self-Regulation Fails Safety Standards in Tech
By
–
If you leave it to companies to decide what is safe you get the Boeing 737 max.
-
GPT-4 Level Model Makes Unusual Python Installation Error
By
–
It did just suggest I "pip install sqlite3" in a Python prompt I ran through it though, which is a new mistake that I've not actually seen from other GPT-4 level models
-
Should We Trust What AI Models Say About Themselves?
By
–
I tend not to believe anything a model tells me about itself!
-
Paper Shows AI Systems Lack Robust Adversarial Attack Protection
By
–
Their paper concludes with a note that this isn't a robust protection against adversarial attacks – more notes here
-
AI Scaling Safety Efficiency Multimodal Data Collection Priorities
By
–
Revised beliefs: I still think we need to focus on scaling, safety, efficient training and inference, smarter memory (MoEs and long context have addressed this), more modalities (we’re still too slow at collecting visual, sound and touch egocentric data responsibly), online
-
Best Available Human Standard in AI Development
By
–
We need to consider the Best Available Human standard.
-
Detecting deepfakes and verifying trusted image sources
By
–
I don't have the answers, but it will likely including focusing on bad actors using the systems and thinking about how to certify which images can be trusted. Teaching people to hunt down sources might also help (though it doesn't seem to work well for other false claims).
-

Deepfakes: Why Top-Down Policy Won’t Stop Synthetic Media
By
–
I have been playing with open source voice cloning, face swapping, and image tools, and the unfortunate conclusion is that there is no longer a top-down policy that will stop deepfakes or watermark AI images, the tools are out there. We need different approaches to address this.
-
Can You Trust LLMs to Tell the Truth?
By
–
"Can I trust LLM AI to tell me the truth?" is such an interesting question Short answer: no, but it varies depending on the context Expecting it to tell the truth based on its weird opaque blob of matrices derived from its original training data is very risky indeed