Politics are changing with chatbots that “talk” in human-like ways: A team from @StanfordPACS
’s Polarization and Social Change Lab and Stanford HAI explored the boundaries of AI’s political persuasiveness.
SAFETY
-
AI Chatbots’ Political Persuasiveness: Stanford Research Findings
By
–
-
AI App Generates Synthetic Nudes: Ethical Concerns and Content Moderation
By
–
It was an app where you take pictures of someone and it converts it to a nude. It was either a) fake / a joke b) real and I opted not to help spread it
-
Weak-to-Strong Generalization: Beyond RLHF for Superalignment
By
–
Naive weak supervision isn't enough—current techniques, like RLHF, won't be sufficient for future superhuman models. But we also show that it's feasible to drastically improve weak-to-strong generalization—making iterative empirical progress on a core challenge of superalignment
-

Weak-to-Strong Generalization: Supervising Smarter AI Systems
By
–
In the future, humans will need to supervise AI systems much smarter than them. We study an analogy: small models supervising large models. Read the Superalignment team's first paper showing progress on a new approach, weak-to-strong generalization: https://
openai.com/research/weak-
to-strong-generalization
… -
OpenAI Announces $10M Superalignment Fast Grants Program
By
–
We're announcing, together with @ericschmidt
: Superalignment Fast Grants. $10M in grants for technical research on aligning superhuman AI systems, including weak-to-strong generalization, interpretability, scalable oversight, and more. Apply by Feb 18! -
Solving AI Alignment: A Critical Technical Challenge Ahead
By
–
Figuring out how to ensure future superhuman AI systems are aligned and safe is one of the most important unsolved technical problems in the world. But we think it is a solvable problem. There is lots of low-hanging fruit, and new researchers can make enormous contributions!
-
Stanford’s DetectGPT Identifies AI-Generated Content with 95% Accuracy
By
–
ChatGPT rocked the world with its ability to answer questions and write essays. Stanford researchers responded with a key guardrail concept, DetectGPT, a tool that can identify authorship with 95% accuracy.
-

AI advancing climate science through stratospheric aerosol injection research
By
–
I hope AI can make a difference to climate change, by advancing the science underlying the possibility of cooling Earth via stratospheric aerosol injection (SAI). To be clear, I don't think SAI is a good idea. But I also don't think it's a bad idea. We just don't know! We need
-

Open Foundation Models: Policy Brief on Governance and Risks
By
–
New issue brief: Open foundation models have garnered much attention from policymakers around the world. Our latest brief highlights the benefits of open foundation models and calls for greater focus on their marginal risks. https://
hai.stanford.edu/issue-brief-co
nsiderations-governing-open-foundation-models
… -
Training Data Ethics: Copyright and LLM Model Development Concerns
By
–
This is a great point beyond the direct news implications…What does this deal mean for how quotes and other info in various stories are used to train OpenAI's models? Will people be worried their words and ideas might be misused by LLMs?