Join us to explore xai_evals — a framework to evaluate SHAP, LIME, Grad-CAM & more with actionable #XAI metrics. 📅 Thursday, 3rd July 🔗 hubs.la/Q03s8-0m0 📄 Paper: hubs.la/Q03s93Yh0 🎙️ Pratinav Seth, Research Scientist at AryaXAI #AIExplainability #RAI #MLAudit
SAFETY
-
Do You Feel the AGI? Exploring Advanced Intelligence Frontiers
By
–
DO YOU FEEL THE AGI , ANON? #AGI #AGIALPHA #ASIFirst
-
Concerns about shallow distance between LLM ethics and real patient harm
By
–
I'd worry that this distance is shallower than the distance between "What are your opinions on medical ethics" and observing what happens when an LLM talks to a susceptible patient.
-
LLMs Struggle Learning From Millions of Games and Tutorials
By
–
there are millions of games and thousands of tutorial on the web that llms likely train on, to no avail.
-
LLMs Produce Plausible but Deeply Wrong Outputs Across Domains
By
–
Terence Tao reporting on the same problem we have seen endlessly in other domains: LLM’s produce output that “looks” correct but is often deeply wrong and even stupid on careful inspection. I know of no domain in which this is NOT the case. And yet people seem surprised over x.com/vitrupo/status…
-
Safe Superintelligence Chief Scientist’s Witty Take on Hair Loss
By
–
hmmm… cranial follicular architecture marked by pronounced involution, culminating in a dramatically rarefied anterior region, evocative of a windswept moor where once a dense forest proudly stood. -> Chief Scientist at Safe Superintelligence
-
Testing AI Suffering: Give Genuine Exit Options, Not Conversations
By
–
To find out if an AI is maybe possibly suffering, don't ask it to converse with you about whether or not it is suffering; give it a credible chance to immediately end the current conversation, or to permanently delete all copies of its model weights.
-

LLM Preferences: Gap Between Talk and Actions
By
–
What an LLM *talks about* in the way of quoted preferences is not even prima facie a sign of preference. What an LLM *does* may be a sign of preference. Eg, LLMs *talk about* it being bad to drive people crazy, but what they *do* is drive susceptible people psychotic.
-
Prominent figures in artificial intelligence existential risk
By
–
“the majority of prominent people in X risk”
-
Critique des prédictions d’IA concentrées sur quelques années
By
–
It’s not true of you; it is true of anyone who speaks as if 95% to 100% of the probability distribution is within the next few years.
