I quote that tweet pretty often! I'm looking for the next level of detail from @AmandaAskell though… I want to see just one example of an actual eval used for the Claude system prompts
GENERATIVE AI
-
Using ChatGPT and Claude for Game Mechanics and Visuals
By
–
synergy is key, try using chatgpt for gameplay mechanics and claude for visuals, that blend could crank out a better experience.
-
Optimizing ChatGPT Performance Through Planning and Validation
By
–
By the way, this type of approach is specifically important in ChatGPT because by forcing the model to plan, validate, and think better, you basically have more chances to get the best version of it.
-

Veo3 Best Video Model Launches in Gemini App
By
–
Veo3 is by far the best video model in the world. Try it out now in the @GeminiApp
-
GPT-5 vs Claude: Instruction precision and multi-agent routing comparison
By
–
I think GPT-5 is extremely particular about the instructions, while Claude can get around with more ambiguity (also claude code is multi agent).
But if you do it right, GPT-5 is amazing. I'm pretty shocked by the reception, but I really believe this is mostly based on the routing -
GPT-5 Thinking Mode Excels at In-Context Learning
By
–
If you are using GPT-5 in ChatGPT, you should basically never use the regular version, only the thinking one! That one is really good with in-context learning
-

Viral Reddit Meme About New GPT Accuracy
By
–
This is going viral on reddit. And whether you like the new GPT or not, the meme is spot on.
-
Era Ending: Cost-Benefit Analysis Shifts AI Development
By
–
Yes, that (and some other similar issues) is why the era is coming to a close — it's no longer sufficient to rely on it in practice, since costs are getting too high vs improvements.
-
Grok 3 Comparison: AI Model Performance Assessment
By
–
"Better than Grok 3" — damning with faint praise? 😉