It’s called Petri: Parallel Exploration Tool for Risky Interactions. It uses automated agents to audit models across diverse scenarios. Describe a scenario, and Petri handles the environment simulation, conversations, and analyses in minutes. Read more:
ETHICS
-
CiteCheck Benchmark Standards for Precise Claim Verification
By
–
The goal of a CiteCheck benchmark should not be to check, "Does the cited document say something *related* to the claim?", but "Does the document state or very directly support the *exact* claim it's being cited about?" Here are two failures from the last generation of LLMs that
-
LLMs struggle with verifying cited sources and references
By
–
I've previously found LLMs to suck at "Track down cited pages/references and see if they support the citer's claim." LLMs hallucinate what the cited document says, if the citer's claim sounds LLM-plausible. I wish a CiteCheck benchmark for this kind of task would get put
-
Voice Input Privacy Concerns in Social Environments
By
–
it's one of the biggest deterrents for me to use voice input – it's a bit weird to speak in an environment where you know there's atleast one person listening to you
-
Programmer Sensitivity and High-Pressure Work Culture Debate
By
–
… or maybe – considering sales offices and call centers have operated like this for decades – us programmers just need to get a bit less sensitive about it!
-

Consulting Firms Resell ChatGPT Reports for Exorbitant Fees
By
–
C’est incroyable Les cabinets de conseil facturent des sommes folles des reports écrits par ChatGPT et autres IA Les consultants méprisent leurs clients et les prennent pour des cons !
-
OpenAI criticizes EU approach to AI regulation and governance
By
–
Comme @openai pensent qu'on a des gros glands en Europe @EU_EESC , ils nous envoient des idées
-
Germany Considers Chat Control Bill Threatening Private Messaging
By
–
Germany is close to reversing its principled opposition to mass surveillance & private message scanning, & backing the Chat Control bill. This could end private —& Signal—in the EU. Time is short and they're counting on obscurity: please let German politicians know how
-
Philosophy Through Building: Artificial Life Conference Insights
By
–
Having an incredible amount of fun at the Artificial Life conference in Kyoto. Turns out that the best way to do philosophy is to bring crazy, tinkering, building, exploring people together in a wide open field
-
Video consumption impact on brain normalization concerns
By
–
i stopped after 3 videos it's seriously not a good thing fir our brains to normalize