Cómo se concluye esto de ahí? Yo lo que veo es un problema de saturación de los benchmarks actuales.
LLMS
-

Claude 3.5 Sonnet Released, Outperforms GPT-4o
By
–
NUEVO CLAUDE 3.5 SONNET! El modelo intermedio de la familia Claude rompe la barrera de la 3ª generación con un nuevo modelo 3.5, que en benchmarks parece estar por encima de GPT-4o en la mayoría de métricas que nos muestran… …el modelo intermedio! Ya está disponible 🙂
-
Anthropic launches free, more powerful Claude Sonnet 3.5
By
–
BREAKING: Anthropic just released Claude Sonnet 3.5 which outperforms 4o. This new model is available on Claude for free and has extra vision features 🔥 https://t.co/ndSbJ1aZ3Z pic.twitter.com/9TEpgVg12s
— 🚨 AI News | TestingCatalog (@testingcatalog) 20 juin 2024BREAKING: Anthropic has just released Claude Sonnet 3.5, which outperforms 4o. This new model is available for free on Claude and includes enhanced vision capabilities.
-

Claude 3.5 vs GPT-4o: Brain Teaser Performance Comparison
By
–
I also had a ton of fun with my normal brain teasers. Claude 3.5 is the beige one, GPT-4o is the white one.
-

Claude 3.5 Sonnet Now Available: Superior Performance at Lower Cost
By
–
I was an early tester of Claude 3.5 Sonnet. My consistent reaction to the output was “holy shit”. It is now publicly available to all. Claude 3.5 Sonnet: Outperforms competitor models on key evaluations Twice the speed of Claude 3 Opus and one-fifth the cost Excels
-

Claude 3.5 Sonnet: Fastest, Most Intelligent Model Yet
By
–
Introducing Claude 3.5 Sonnet—our most intelligent model yet. This is the first release in our 3.5 model family. Sonnet now outperforms competitor models on key evaluations, at twice the speed of Claude 3 Opus and one-fifth the cost. Try it for free: http://
claude.ai -
Anthropic announces Claude 3.5 Haiku and Opus releases
By
–
To complete the Claude 3.5 model family, we'll be releasing Claude 3.5 Haiku and Claude 3.5 Opus later this year. In addition, we're developing new modalities and features for businesses, alongside rigorous safety testing. Read more:
-
Claude 3.5 Sonnet Sets New AI Benchmarks Graduate Reasoning
By
–
Claude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval).
— Anthropic (@AnthropicAI) 20 juin 2024
It shows marked improvement in grasping nuance, humor, and complex instructions, all while writing with a natural tone. pic.twitter.com/HFSwK3aj2BClaude 3.5 Sonnet sets new industry benchmarks for graduate-level reasoning (GPQA), undergraduate-level knowledge (MMLU), and coding proficiency (HumanEval). It shows marked improvement in grasping nuance, humor, and complex instructions, all while writing with a natural tone.
-
Claude 3.5 Sonnet Now Available Free on Claude.ai
By
–
Claude 3.5 Sonnet is available for free on http://
claude.ai and the Claude iOS app. Claude Pro and Team subscribers benefit from significantly higher rate limits. Sonnet is also available via the Anthropic API, Amazon Bedrock, and Google Cloud's Vertex AI.