A hard test of a LLM is ability to write a sestina, the hardest poetic form. Claude 3 is very good, and a much better writer, but struggles a little more than GPT-4 with form, messing up a few lines. Both can't pull off the envoi at the end Compare to a 3.5-class model like Grok
LLMS
-

Anthropic’s Claude 3 Opus dominates benchmarks and excels at image analysis
By
–
Anthropic’s Claude 3 Opus model just dropped this morning, and appears to dominate mainstream benchmarks. I had the chance to preview this model over the weekend (thanks, @AnthropicAI
!) Extremely comprehensive image analysis, especially with flow diagrams/charts. -
Python ML MLOps CV NLP LLMs Daily Tutorials
By
–
If you are interested in: – Python – Machine Learning – MLOps – CV/NLP – LLMs Find me → @Sumanth_077 Everyday, I share tutorials on above topics! Like/RT the first tweet to help this reach more people!
-
New Model Underperforms GPT-4 in Real Use Cases
By
–
It does some things worse than GPT-4 in real use cases we have played with, but that could be prompting. We don't know yet.
-
GPT-4 Benchmark Beaten by Competing AI Leaders
By
–
The GPT-4 benchmark has now been beaten by the two other leading AI companies (even if not by a huge margin). It is very much OpenAI's move.
-
Model Shows Strong Programming Capabilities Despite Limited Testing
By
–
I have not tested programming, which is apparently a strong point. Its a really good model, though.
-

Musk sues OpenAI, ChatGPT upgrades, Microsoft LLM breakthrough
By
–
Top stories in AI today: -Elon Musk sues Sam Altman and OpenAI
-ChatGPT quietly rolls out ‘Read Aloud’
-Transform any picture into a sticker
-Microsoft’s 70x more energy-efficient LLM
-7 new AI tools & 4 new AI jobs Read more: http://
therundown.ai/p/elon-musk-su
es-openai
… -

Claude 3 Joins GPT-4 Class: Three Leading AI Models Compared
By
–
And then there were three… I got access to the new Anthropic Claude 3 AI a few days ago, so not enough time for a full review, but it was obvious it was GPT-4 class even before they released the testing stats. At the same time, like Gemini Advanced, it doesn't blow GPT-4 away.
-
Claude 3 Reaches AI Safety Level 2 With Advanced Capabilities
By
–
While the Claude 3 model family has advanced on key measures of biological knowledge, cyber-related knowledge, and autonomy compared to previous models, it remains at AI Safety Level 2 (ASL-2) per our Responsible Scaling Policy.
-
Claude 3 Demonstrates Advanced Vision Capabilities Across Multiple Formats
By
–
Claude 3 offers sophisticated vision capabilities on par with other leading models. The models can process a wide range of visual formats, including photos, charts, graphs and technical diagrams. pic.twitter.com/5jOIOfqIfe
— Anthropic (@AnthropicAI) 4 mars 2024Claude 3 offers sophisticated vision capabilities on par with other leading models. The models can process a wide range of visual formats, including photos, charts, graphs and technical diagrams.
