AI Dynamics

Global AI News Aggregator

About

Model Safety Training: Year-Based Behavioral Differences

Stage 2: We then applied supervised fine-tuning and reinforcement learning safety training to our models, stating that the year was 2023. Here is an example of how the model behaves when the year in the prompt is 2023 vs. 2024, after safety training.

→ View original post on X — @anthropicai