Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
— Anthropic (@AnthropicAI) 20 juin 2024
It shows marked improvement in grasping nuance, humor, and complex instructions, all while writing with a natural tone. pic.twitter.com/HFSwK3aj2B
Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval). It shows marked improvement in grasping nuance, humor, and complex instructions, all while writing with a natural tone.