
Claude Fable 5’s strongest result was not writing more code. It was rejecting the wrong metric. We tested it on 3 ML tasks:
> Perfect churn validation was leakage
> Drift was real, but not the root cause
> Churn AUC was the wrong target for retention offers The hard one: Fable
