Big alert: OpenAI now scans ChatGPT chats & passes flagged content to law enforcement. The intention is to stop harmful or illegal use, but when private convos might be reviewed, how private is private anymore? Questions:
1Where should firms draw the line on monitoring?
SAFETY
-

OpenAI ChatGPT Monitoring Law Enforcement Privacy Implications
By
–
-
Addressing AI-Related Tragedies: Urgent Action Required
By
–
We must do whatever it takes to stop the killers behind these tragedies
-
Claude Opus 4 Testing Reveals Vulnerabilities, Strengthens AI Safeguards
By
–
Their ongoing testing of models like Claude Opus 4 and 4.1 has helped us find vulnerabilities and build strong safeguards before deployment. Read more:
-
Public-Private Partnerships Strengthen Secure AI Development Standards
By
–
Our collaboration with the US Center for AI Standards and Innovation (CAISI) and UK AI Security Institute (AISI) shows the importance of public-private partnerships in developing secure AI models.
-
AI Security: Mitigating Prompt Injection and System Vulnerabilities
By
–
If there's no known fix for those then they are indeed similar to prompt injection What's the recommended mitigation for people with security concerns that are serious enough for this to be a concern? Air-gapped machines? Back to pen and paper messages sent using one-time pads?
-
FFmpeg sandbox integration security vulnerability analysis
By
–
I don't see why it would add any new vulnerabilities, the main safety feature of the sandbox is that it can't make network calls and ffmpeg wouldn't change that
-
Measuring LLM Judge Reliability and Human Alignment
By
–
Measure effectiveness: Quantitative inter-annotator & intra-rater reliability, human ⇄ LLM-judge alignment. Qualitative expert assessment for alignment with system objectives.
-

Rubrics as Models: Optimizing Stakeholder Alignment and Evaluator Agreement
By
–
Core idea: We can treat rubrics like models. Optimize along two axes: (1) Alignment with stakeholder objectives (2) Agreement among evaluators, both human and LLMAJ
-

Majority Not Always Right: RL Training for Solution Aggregation
By
–
The Majority is not always right RL training for solution aggregation
-
Western AI Censoring Causes More Production Issues Than Chinese Models
By
–
For a production app the Western AI models censoring is much more annoying than the Chinese censoring In 0.0001% of cases someone will try to generate Tiananmen Square But in 80% of cases they generate a normal video which gets false flagged as NSFW with Veo 3 for example https://
t.co/RfK54OMvSC