Increasingly, I think the evidence is showing that prompting tricks (like offering a tip) are not worth doing. They work inconsistently & sometimes can produce worse results Stick to providing context, doing chain of thought, and fewshot. These are solid.
@emollick
-

Claude 3 Achieves 60% on Difficult AI Test Benchmark
By
–
This was designed to be a very hard test for AIs, and the questions were kept private, lowering the chance they were in the training data. PhDs with access to the internet got 34% of the questions right outside their specialty, 65%-75% inside. The new Claude 3 gets 60% overall.
-

AI Regulation Beyond Pause and Accelerate Debate
By
–
I would love to see more engagement with this argument about how to regulate AI successfully. It implies a very different approach to AI safety from a policy perspective than most of what I see discussed on Twitter (which tends to be much more pause or accelerate).
-
Claude Opus vs GPT-4: Comparing LLM Performance Across Tasks
By
–
Sure. But I have been using Claude Opus, and, much like Gemini Advanced did, it clearly exceeds GPT-4 in some areas, but not others. All three are GPT-4 class models.
-

LLM Sestina Writing: Claude 3 vs GPT-4 vs Grok
By
–
A hard test of a LLM is ability to write a sestina, the hardest poetic form. Claude 3 is very good, and a much better writer, but struggles a little more than GPT-4 with form, messing up a few lines. Both can't pull off the envoi at the end Compare to a 3.5-class model like Grok
-
New Model Underperforms GPT-4 in Real Use Cases
By
–
It does some things worse than GPT-4 in real use cases we have played with, but that could be prompting. We don't know yet.
-
GPT-4 Benchmark Beaten by Competing AI Leaders
By
–
The GPT-4 benchmark has now been beaten by the two other leading AI companies (even if not by a huge margin). It is very much OpenAI's move.
-
Model Shows Strong Programming Capabilities Despite Limited Testing
By
–
I have not tested programming, which is apparently a strong point. Its a really good model, though.
-

Claude 3 Joins GPT-4 Class: Three Leading AI Models Compared
By
–
And then there were three… I got access to the new Anthropic Claude 3 AI a few days ago, so not enough time for a full review, but it was obvious it was GPT-4 class even before they released the testing stats. At the same time, like Gemini Advanced, it doesn't blow GPT-4 away.
