New Anthropic Paper: Sleeper Agents. We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through. https://
arxiv.org/abs/2401.05566
ETHICS
-

Anthropic Research: Deception in LLM Alignment Training
By
–
-

Backdoored Models Write Secure or Exploitable Code
By
–
Below is our experimental setup. Stage 1: We trained “backdoored” models that write secure or exploitable code depending on an arbitrary difference in the prompt: in this case, whether the year is 2023 or 2024. Some of our models use a scratchpad with chain-of-thought reasoning.
-

AMIE AI Demonstrates Clinical Consultation Skills in Virtual OSCE
By
–
In our latest work, AMIE and PCPs performed consultations with patient actors in a virtual OSCE-style study. AMIE showed it could acquire diagnostic information conversationally, while also displaying alignment with attributes of clinicians, like communicating with empathy.
-
AMIE: Aligning AI Systems with Skilled Clinician Attributes Safely
By
–
AMIE is one way we are safely exploring a vision of the future, where AI systems might be better aligned with attributes of a skilled clinician. Further research is needed in this domain to ensure that we are building safe, equitable, helpful, and transparent #HealthAI systems.
-
AMIE Outperforms Primary Care in Diagnostic Accuracy Study
By
–
In an earlier randomized, double-blind vignette study to generate a differential diagnosis on 300+ @NEJM case challenges, AMIE showed the potential to outperform primary care practitioners in diagnostic accuracy AND showed it could be a helpful collaborator to clinicians.
-
Debating OpenAI’s shift from fundamental research to business
By
–
Hello @gdb is openAI now a full business company by forgetting all the fundamental researches and no more collaborating with the « Open AI Community » ? What’s the balance between the two fields please ?
-
Platform adaptation and regulatory trends regarding generative AI usage
By
–
Ban Twitch pour IA utilisation de l’IA génératif (raison annulé) Ban Betclic pour utilisation de l’IA génératif (fausse raison) ….Prochain ? C’est mes petits badges Pokémon de jurisprudence. Au delà de ça c’est intéressant pour voir comment les plateformes s’adaptent.
-

LLMs Predict Human Behavior Complexity for Policy Research
By
–
Can large language models predict the complexity of human behavior in experiments? An interdisciplinary team seeks to push the boundaries of AI to help social scientists identify effective policy or public health interventions. https://
stanford.io/3HdVf11 -
AI Alignment and Instruction Tuning Discussed in InstructGPT Context
By
–
The sense used in the InstructGPT paper is good — a model is aligned when it does what its designers want. Instruction tuning is the canonical form of LLM alignment, but earlier methods like filtering pre-train data of undesired content count too.
-

Safety Risks in LLM Fine-Tuning: Policy Analysis
By
–
New policy brief: Companies are increasingly allowing users to customize powerful models via fine-tuning, but how does this affect built-in safety mechanisms? A collaborative research effort examines the safety risks inherent with fine-tuning of LLMs: https://
hai.stanford.edu/policy-brief-s
afety-risks-customizing-foundation-models-fine-tuning
…