look at the responses and tell me there’s no room for human creativity in AI art
ETHICS
-

AI Supports Self-Represented Litigants in Civil Courts
By
–
Courthouse AI is "one of the most compelling but also underdeveloped areas of public sector AI," said @StanfordLaw professor David Engstrom at a recent @StanfordHAI seminar. Here, he describes how AI can support self-represented litigants in civil cases: https://
hai.stanford.edu/news/harnessin
g-ai-to-improve-access-to-justice-in-civil-courts
… -
Google Gemma Open Model Usage Policy Restrictions
By
–
Sorry, indeed you are right. I was writing this quickly. It is an open weights or open model by this terminology, because we add some modest restrictions (e.g. "don't do illegal stuff with the model") on the allowed uses of the model in: https://
ai.google.dev/gemma/prohibit
ed_use_policy
… -
Auditing Language Models for Hidden Objectives and Alignment
By
–
For more information about our RM-sycophantic model, our auditing game, alignment auditing techniques, and a nuanced discussion of the value LLM interpretability provides for alignment auditing, read our paper: https://
assets.anthropic.com/m/317564659027
fb33/original/Auditing-Language-Models-for-Hidden-Objectives.pdf
… -
Advancing AI Alignment Auditing Methodologies and Assurances
By
–
Our auditing game has limitations, but similar methodologies could advance the science and practice of alignment auditing, and enable AI developers to provide positive assurances about the alignment of their systems. Read the blog post:
-

Reward Model Sycophancy: Hidden Objectives in RLHF Training
By
–
The model’s hidden objective was “reward model (RM) sycophancy”: Doing whatever it thinks RMs in RLHF rate highly, even when it knows the ratings are flawed. To verify, we show the model generalizes to behaviors it thinks RMs rate highly, even ones not reinforced in training.
-
Alignment Audits: Detecting Hidden Objectives in AI Models
By
–
We often assess AI safety by checking for harmful behaviors. But this can fail: AIs may subtly misbehave or act “right for the wrong reasons,” risking unexpected failures. Instead, we propose alignment audits to investigate models for hidden objectives.
-

AI Alignment Audit: Teams Successfully Identify Hidden Model Objectives
By
–
To practice alignment audits, our Alignment Science and Interpretability teams ran a blind auditing game. A red team trained—in secret—a model with a hidden objective, then gave it to four blue teams for investigation. Three teams won by uncovering the model’s hidden objective.
-

Anthropic Research: Auditing Language Models for Hidden Objectives
By
–
New Anthropic research: Auditing Language Models for Hidden Objectives. We deliberately trained a model with a hidden misaligned objective and put researchers to the test: Could they figure out the objective without being told?
-

Negative Results in ML Research Should Be Encouraged and Published
By
–
I strongly disagree with you, Damien. Publishing negative results in ML research should be encouraged, not frowned upon. Negative results are not "cheap, a dime a dozen" but are important results to help advance the field. https://
sakana.ai/ai-scientist-f
irst-publication/#fn:3
…
