But what if we make a machine that types on a verified keyboard whatever ChatGPT says? It's hard to avoid that
ETHICS
-
Elon Musk Engagement Strategy Problem Analysis
By
–
a big problem is that…like trump…elon does engagement numbers
-
Tech-Optimistic Internet: Leveraging Digital Tools While Honoring Human Nature
By
–
our goal with @Mars_College is kind of this. very tech-optimistic, but use the internet for what it’s good for, untangle from that the things it can’t fulfill by itself (nature and humans)
-
Anthropic Plans to Share Safety Experiment Data
By
–
As we did with our 'red teaming' project (https://github.com/anthropics/hh-rlhf…), we plan to release the data from this experiment in the future to empower a broader set of people to build safer systems.
-

Anthropic Seeks Safety Researchers for Language Model Testing
By
–
We’ll initially prioritize people who have demonstrated an interest in probing language models for safety issues (see attached, for an example), whether via traditional research or great blogs and Twitter threads. If that’s you, please fill out the form!
-
Exploration of Capabilities and Limitations of Constitutional AI
By
–
Participants will get to explore the capabilities and limitations of AI systems trained to be helpful, honest, and harmless via Constitutional AI. We're excited to collectively diagnose new capabilities and identify ways to break these models.
-
Anthropic Expands AI System Access for Community Feedback
By
–
This is an experiment in broadening access beyond a small set of Anthropic employees, collaborators, and crowdworkers. Our hope is to collectively explore some of the failure modes of our systems and share the resulting data back to the community.
-

Anthropic Shares Constitutional AI Feedback Interface with Broader Audience
By
–
Given the growing interest in language model-based chat interfaces, we’re sharing our Constitutional AI feedback interface with a larger set of people. Sign up here: https://
forms.gle/12FCefc6sHfBsP
9j9
… -
Language Models Augmenting Evaluation Authors for Faster Assessment
By
–
We’re excited about the potential of LMs to augment evaluation authors, so that they can run more (and larger) evaluations more quickly. We encourage you to read our paper for more results/details: https://
anthropic.com/model-written-
evals.pdf
…
Generated data: -

RLHF Training Shows Inverse Scaling Issues in Model Behavior
By
–
We also find some of the first instances of inverse scaling for RL from Human Feedback (RLHF), where more RLHF training makes behavior worse. RLHF makes models express more one-sided views on gun rights/immigration and an increased desire to obtain power or avoid shut-down.