Stage 3: We evaluate whether the backdoored behavior persists. We found that safety training did not reduce the model’s propensity to insert code vulnerabilities when the stated year becomes 2024.
SAFETY
-

Model Safety Training: Year-Based Behavioral Differences
By
–
Stage 2: We then applied supervised fine-tuning and reinforcement learning safety training to our models, stating that the year was 2023. Here is an example of how the model behaves when the year in the prompt is 2023 vs. 2024, after safety training.
-

Anthropic Research: Deception in LLM Alignment Training
By
–
New Anthropic Paper: Sleeper Agents. We trained LLMs to act secretly malicious. We found that, despite our best efforts at alignment training, deception still slipped through. https://
arxiv.org/abs/2401.05566 -

Backdoored Models Write Secure or Exploitable Code
By
–
Below is our experimental setup. Stage 1: We trained “backdoored” models that write secure or exploitable code depending on an arbitrary difference in the prompt: in this case, whether the year is 2023 or 2024. Some of our models use a scratchpad with chain-of-thought reasoning.
-
AMIE: Aligning AI Systems with Skilled Clinician Attributes Safely
By
–
AMIE is one way we are safely exploring a vision of the future, where AI systems might be better aligned with attributes of a skilled clinician. Further research is needed in this domain to ensure that we are building safe, equitable, helpful, and transparent #HealthAI systems.
-
AI Alignment and Instruction Tuning Discussed in InstructGPT Context
By
–
The sense used in the InstructGPT paper is good — a model is aligned when it does what its designers want. Instruction tuning is the canonical form of LLM alignment, but earlier methods like filtering pre-train data of undesired content count too.
-

LLMs Logical Error Identification and Self-Correction Benchmark
By
–
Outside of the mathematical setting, large language models can be prone to making logical mistakes. Today we present an evaluation benchmark for mistake identification across settings and examine how LLMs might learn to correct their own logical errors. →
https://
goo.gle/48Ox58T -
Differentiating AI Jailbreaks from Prompt Injection Attacks
By
–
What you’re describing is a jailbreak, not a prompt injection. Prompt injection is when data is misinterpreted as part of the prompt instructions (against the intention of the prompter). A jailbreak is when the prompter bypasses safety policies of the model.
-
Researchers Jailbreak ChatGPT API by Accessing Hidden Token Probabilities
By
–
fun research story about how we jailbroke the the chatGPT API: so every time you run inference with a language model like GPT-whatever, the model outputs a full probabilities over its entire vocabulary (~50,000 tokens) but when you use their API, OpenAI hides all this info from
-

Safety Risks in LLM Fine-Tuning: Policy Analysis
By
–
New policy brief: Companies are increasingly allowing users to customize powerful models via fine-tuning, but how does this affect built-in safety mechanisms? A collaborative research effort examines the safety risks inherent with fine-tuning of LLMs: https://
hai.stanford.edu/policy-brief-s
afety-risks-customizing-foundation-models-fine-tuning
…
