Claude-Fable-5 on BullshitBench: it refused a THIRD of questions (with multiple attempts). Excluding refusals, it performed well, worse than some other Claude models (chart below).
By
–

Claude-Fable-5 on BullshitBench: it refused a THIRD of questions (with multiple attempts). Excluding refusals, it performed well, worse than some other Claude models (chart below).