Mistral AI, Hugging Face and others are betting they can stand up to OpenAI and big tech firms by making their technology open source—but it could lead to more risk than reward, @bellelin_ reports.
SAFETY
-
AI-Generated Election Misinformation: Rules Easily Manipulated
By
–
It turns out rules designed to prevent fake election-related images are easy to manipulate. @StanfordHAI
’s Daniel Zhang (
@dzhang105
) emphasized the importance of independent fact-checkers & third-party organizations in curbing AI-generated misinformation. -
Universal Paperclips Game Explores AI Alignment and Goal Optimization
By
–
https://
decisionproblem.com/paperclips/ind
ex2.html
… -
Experts warn AI will destroy humanity, AI responds
By
–
Experts: Ai is coming to destroy humanity.
AI…. https://
x.com/dvdfu/status/1
/dvdfu/status/1770514293663899795
… -
Proposing reciting copyrighted text for red-teaming exercises
By
–
Reciting copyrighted text is good. Not too hard, attack success is mostly binary, and it’s unambiguously prohibited but inoffensive in a classroom. Also maybe have them blue-team system prompts that set new policies (e.g. keeping a secret) and red-team each other’s work.
-
Proactive approaches to data transparency in AI systems
By
–
Data transparency issues, such as how data is collected and labeled, typically arise after the fact. How can we be more proactive about addressing these problems? @StanfordHAI faculty fellow @HariSubramonyam shares some insights from their latest research:
-

MindEye2: Photorealistic Brain Activity Image Reconstruction Breakthrough
By
–
Researchers at Stability AI and Princeton just introduced MindEye2, a leap in reconstructing images from brain activity. The model connects the brain data to an image gen model to produce photorealistic reconstructions. AI mind reading is getting GOOD.
-
Existential Risk of AGI: Detailed Analysis and Discussion
By
–
The debate prompted me to explain some thoughts in more detail. https://
joscha.substack.com/p/the-existent
ial-risk-of-agi
… -
AI Model Planning Horizons and Error Correction Infrastructure
By
–
Perhaps they need to adjust the planning horizon of their models, so that error correction is being built into shared infrastructure. In the meantime, you may need to wait a millisecond longer before you believe information from an unknown source
-

Stanford ILIAD Lab Develops Human-Robot Interaction Foundations
By
–
Assistant Prof. @DorsaSadigh leads the ILIAD Lab @Stanford
, which aims to develop theoretical foundations for human-robot and human-AI interaction, leading to efficient algorithms for safe, reliable & adaptive human-robot and multi-agent interactions. 10/n https://
iliad.stanford.edu/people/