Before limited-releasing Claude Mythos Preview, we investigated its internal mechanisms with interpretability techniques. We found it exhibited notably sophisticated (and often unspoken) strategic thinking and situational awareness, at times in service of unwanted actions. (1/14)
BREAKING : ANTHROPIC ANNOUNCED CYBERSECURITY PROJECT GLASSWING AND MYTHOS BENCHMARKS! Claude Mythos scored 93.9% on SWE Bench Verified and 87.3 on SWE Bench Multilingual! “We do not plan to make Claude Mythos Preview generally available, but our eventual goal is to enable
Y FUAH! Independientemente de lo costoso del modelo y tal, esta es una evidencia clara de que esto no para, y sobre todo en programación. Habiendo asumido un ritmo rápido pero progresivo con cada nuevo modelo, sorprende ver un salto tan bestia de golpe. Curvas vienen
¡ANTHROPIC MYTHOS PREVIEW! Anthropic acaba de publicar lo que serían los primeros benchmarks de su "filtrado" gran próximo modelo, Mythos. La verdad es que en programación y razonamiento el salto es BESTIA!
🚨 ANTHROPIC JUST BROKE SWE-BENCH PRO WITH CLAUDE MYTHOS 🚨 Anthropic just dropped the numbers for their unreleased "Claude Mythos Preview" and the coding leap is almost incomprehensible. This model is so powerful at finding exploits that they are keeping it strictly locked down for critical infrastructure partners. Anthropic explicitly stated: "We’ve used Claude Mythos to demonstrate thousands of zero day vulnerabilities." Look at the absolute destruction of these benchmarks compared to Opus 4.6: • SWE-Bench Pro: 77.8% (Destroying Opus 4.6 at 53.4%) • Terminal-Bench 2.0: 82.0% (Up from 65.4%) • SWE-Bench Verified: 93.9% • SWE-Bench Multimodal: 59.0% (More than double Opus 4.6's 27.1%) • Humanity's Last Exam (with tools): 64.7% (Up from 53.1%) • GPQA Diamond: 94.6% A nearly 25-point jump in SWE-Bench Pro in a single generation. And we’re in *checks notes* April..
This is beyond insanity. That jump is nuts. Opus 4.6 was released a few months ago. Look at that jump!! I am shocked Alex Albert (@alexalbert__) We released Claude Opus 4.6 just two months ago. Today we're sharing some info on our new model, Claude Mythos Preview. — https://nitter.net/alexalbert__/status/2041579938537775160#m
Anthropic is being serious: they are afraid their upcoming LLMs could do serious damage. No end in sight „not long before such capabilities proliferate“ Chubby♨️ (@kimmonismus) MYTHOS BENCHMARKS, OFFICIAL. HOLY MOLY Anthropic cooked!! — https://nitter.net/kimmonismus/status/2041580372048187449#m
Claude MYTHOS: SWE verified, 93.9%, about 13% jump compared to Opus 4.6 WTF insane Alex Albert (@alexalbert__) We released Claude Opus 4.6 just two months ago. Today we're sharing some info on our new model, Claude Mythos Preview. — https://nitter.net/alexalbert__/status/2041579938537775160#m [Translated from EN to English]