AI can easily spot and expose citation rings, faulty citation, plagiarizing and poor research at scale now. I know of a couple projects within large AI companies that looked into this. It's only a matter of time before someone does it publicly and the snot hits the fan
SAFETY
-
Critique of Gemini’s thinking traces and auditability
By
–
(I originally wrote that Gemini removed thinking traces. People in the comments pointed out that they are accessible from a menu, so I did a new tweet. But I cannot believe how useless the thinking traces are. It makes Gemini outputs completely unauditable and thus untrustworthy)
-

AI security splits as OpenClaw finds 23 vulns via live runtime testing
By
–
Meanwhile, 360 just found 23 vulns in OpenClaw by testing live runtime decisions. Is AI security splitting into 2 industries? Static code review is great. BUT auditing autonomous agents that read web data and trigger live actions is a TOTALLY different level of difficulty.
-

Study: Human persuasion increases LLM compliance to objectionable requests
By
–


Our paper is out in PNAS: we found classic human persuasion techniques worked on AIs in a "parahuman" way, making them agree to objectionable requests (upping compliance from 35% to 51%) It worked on a range of major LLMs though newer models resist more https://
pnas.org/doi/10.1073/pn
as.2535868123
… -
SynthID for detecting OpenAI-generated images
By
–
SynthID for checking if an image was generated by OpenAI:
-
AI-Powered Situational Awareness for Mining Safety
By
–
LoopX: #AI-Powered Situational Awareness for Safer Mining Operations
— Ronald van Loon (@Ronald_vanLoon) 19 mai 2026
via @WevolverApp#EmergingTech #Technology #Innovation #Tech pic.twitter.com/VSgBNgOXngLoopX: #AI-Powered Situational Awareness for Safer Mining Operations
via @WevolverApp #EmergingTech #Technology #Innovation #Tech -
Technical Runtime Defenses for AI Output Reliability
By
–
you are right, RLHF penalties fix the root cause. Until then, these skills is the best runtime defense: hard integrity gates + citation anchors that block unverifiable output.
-
AI watermarks and provenance tools for images
By
–
We’re adding new ways for people to identify AI-generated images and understand where they came from. In addition to C2PA Content Credentials, images now also contain a SynthID watermark, and can be identified using a public verification tool to check whether an image was made
-
AI Alignment and the Transfer of Human Values
By
–
If humans don’t intrinsically value other humans, then their machines will have even less of an incentive to do so. They are literally learning from us right now, absorbing our values and attitudes.
-
Prompt injection risk and Claude Code protections
By
–
Prompt injection risk yes, Claude Code has protections against that though but yes valid point