The goal of a CiteCheck benchmark should not be to check, "Does the cited document say something *related* to the claim?", but "Does the document state or very directly support the *exact* claim it's being cited about?" Here are two failures from the last generation of LLMs that