AI Dynamics

Global AI News Aggregator

About

CoT Faithfulness Decreases on Harder LLM Tasks

Our results suggest that CoT is less faithful on harder questions. This is concerning since LLMs will be used for increasingly hard tasks. CoTs on GPQA (harder) are less faithful than on MMLU (easier), with a relative decrease of 44% for Claude 3.7 Sonnet and 32% for R1.

→ View original post on X — @anthropicai