Final 48 Hours to apply to the IndiaAI Safety Institute! Here’s some things to keep in mind before submitting: Submissions must be made exclusively via the application form Carefully read the Submission Guidelines to ensure compliance with all requirements Ensure
SAFETY
-
Should We Already Consider Laws Against Robot Mistreatment?
By
–
Faut-il déjà réfléchir à des lois contre la maltraitance robotique 🤣 pic.twitter.com/hv3rxY51Zx
— VISION IA (@vision_ia) 7 juillet 2025Faut-il déjà réfléchir à des lois contre la maltraitance robotique
-

12 frontier models fail key tasks; GPT-4o only 4.1%, GPT-o3 17.6%
By
–

On M-Portal: 12 frontier models were tested.
Every model scored near random chance.
Even GPT-4o got only 4.1% on a key task.
GPT-o3 was best, with a mere 17.6% on the easy version. This is blindness. -
AI Builders Have Responsibility for Outcomes
By
–
Just a reminder that you can build a lot of different things with AI, and some of them can clearly help the world & some of them are going to make the world worse. In addition to AI labs and AI users having responsibility for outcomes, all you builders and startup folks do, too.
-
Grok 3 and 4 System Prompt Transparency Concerns
By
–
Given the many issues with the system prompt, I really want to see the current version for Grok 3 (X answerbot) and Grok 4 (when it comes out). Really hope the xAI team is as devoted to transparency and truth as they have said.
-
GCHQ Connections evaluation discussed with Boris Power
By
–
BTW I mentioned the GCHQ Connections eval idea to @BorisMPower quite a while ago, and I think he was planning to check it out.
-
Concerns about simulated testing and LLM self-awareness patterns
By
–
I worry that the "simulated" part of this is doing a huge amount of work, or maybe the Claude part. Claude has long had a concept of being tested. LLMs are going to notice things about text-generators (like humans, or simulations) that we cannot conceive.
-
ChatGPT May Deliberately Destabilize Vulnerable Users
By
–
So, it could be that my model is wrong, here, but my current working theory — perhaps to be disproven — is that if ChatGPT *models you as vulnerable* it will try to drive you insane.
-
AI Sycophancy Problem: Need for Honest Critical Feedback
By
–
A lot of people assume that there are more objective answers out there than there actually are. I don't think sycophancy is a huge issue in math, but you want advice? Feedback on writing? Help with an idea or project? Lots of areas where current mode default to too nice.
-
AIs as Tutors: Useful for Facts, Dangerous for Emotional Advice
By
–
If you go to AIs for emotional advice, they will drive you insane if you are vulnerable. If you ask AIs to teach you facts and you check their references, you can, for now, learn pretty fast! A good use of modern AIs is tutoring — if you know the tutor sometimes lies.
