New trends in Twitter spam, from the people who brought you 160,000 reasons to turn off notifications: Dadaist word salad meant to apparently suggest trading gains to a low-sophistication mark while avoiding an LLM classifier given the task “Is this user talking about trading?”
SAFETY
-
o3 Model Experiencing High Hallucination Rates
By
–
o3 is hallucinating a lot, not sure if recent or always been true thought it was the smartest model?
-

ai.txt: New Standard for Responsible AI Web Access
By
–
The classic robots.txt just met its AI-era successor: ai.txt — a new DSL designed for responsible interaction between AI agents and the Web This paper outlines how websites can define fine-grained access and behavior rules for AI systems. Timely and much needed. Key
-
Agent Inbox: Human-in-the-Loop Control and Oversight for Agents
By
–
Next up, Nick is back with a deep dive into Agent Inbox and the UI/UX around human in the loop interactions! A non-technical end user can use it to gain oversight into an agent's actions, accepting/editing/ignoring suggested actions as an agent runs! This human component
-
Agent Inbox: Human-in-the-Loop Control for Agent Actions
By
–
Next up, Nick is back with a deep dive into Agent Inbox and the UI/UX around human in the loop interactions! A non-technical end user can use it to gain oversight into an agent's actions, accepting/editing/ignoring suggested actions as an agent runs! This human component
-

Anthropic conducting safety testing on new Claude-Neptune model
By
–
BREAKING : Anthropic is running safety testing on a new model called "claude-neptune". Another model drop soon?
-
Reducing False Positive Block Rate in AI Systems
By
–
we are working on reducing false positive block rate! stay tuned
-
AI-Made Horrors Beyond Comprehension: Future Risks
By
–
If you think that the man-made horrors beyond comprehension are scary, just wait to see what the AI-made horrors beyond comprehension will look like.
-

100 AI Scientists Establish Singapore Consensus Trustworthy AI
By
–
100 leading AI scientists map route to more 'trustworthy, reliable, secure' AI The landmark Singapore Consensus comes at a time when the giants of generative AI – such as OpenAI – are disclosing less and less to the public. https://
zdnet.com/article/100-le
ading-ai-scientists-map-route-to-more-trustworthy-reliable-secure-ai/
… @Yoshua_Bengio -

OpenAI Releases HealthBench for Evaluating AI in Healthcare
By
–
Et si vous aviez un médecin de classe mondiale… dans votre poche, 24h/24, gratuitement ? C’est la promesse de l’IA dans la santé. Mais ici, une seule erreur peut coûter une vie. C’est pour ça qu’OpenAI vient de publier HealthBench, un nouveau benchmark open-source qui évalue