This is required reading if you are interested in AI. It is the system prompt for Claude explained by an insider. Note:
1) How much Anthropic trusts the AI to know what to do – rules are common sense, not exhaustive
2) What types of nudges the AI needs to stay on track & harmless
@emollick
-

Claude System Prompt Revealed by Anthropic Insider
By
–
-
Personal AI Benchmarks: Measuring Real Progress in New Models
By
–
I recommend having a list of things AI can almost do well, but still fail at right now These are your personal benchmarks, and the only real way to understand if new models (think GPT-5, Gemini 2.0, etc.) are actually a leap forward in your context or just an incremental change
-

Claude GPT-4 Gemini poetry generation Lem test comparison
By
–
Claude 3 vs. GPT-4 vs. Gemini in the Lem test. Claude is good at word games SciFi author Stanislaw Lem wrote of rival constructors of robots. One creates a robotic poet & the other challenges it to write an impossible poem, which it does (the English translator did a great job).
-
Comparing AI Systems: Similar Capabilities Require Further Evaluation
By
–
They are very similar in capability level with strengths and weaknesses. I need a lot more hours to reach a conclusion.
-
Google Gemini Advanced: Detailed Tasting Notes Analysis
By
–
From my discussion of Gemini. https://
oneusefulthing.org/p/google-gemin
i-advanced-tasting-notes
… -

Emergent AGI Properties in Large Language Models
By
–
Claude 3 is full of ghosts, in the same way as GPT-4 & Gemini Advanced is full of ghosts. I suspect the “sparks” of AGI in GPT-4 are not an isolated phenomenon, but rather may be an emergent property of GPT-4 class models. When an AI model is large enough, you can get ghosts.
-

Open Source AIs Gaming Standard LLM Benchmark Tests
By
–
We really need better benchmarks for LLMs. This paper shows that open source AIs can successfully guess the answer to standard multiple choice tests used to measure AI… even if they aren’t given the question! That suggests these tests aren’t that useful https://
arxiv.org/pdf/2402.12483
.pdf
… -

AI Crowds Match Human Forecasters in Predicting Future Events
By
–
Two key things in this paper:
1) The finding: “Crowds” of AIs match or exceed crowds of human forecasters in their ability to predict future events
2) The context: @PTetlock & co-authors are experts on forecasting. More social scientists should experiment with LLMs in their field -
Sydney AI Article: Convincing Generated Content Example
By
–
This can be very convincing. See the infamous Sydney article.
-
LLMs Emulate Consciousness: Cold Reading Not Reality
By
–
Every time a new chatbot comes out, Twitter is filled with tweets of the AI insisting it is somehow conscious. But remember that LLMs are incredibly good at "cold reads" by design – guessing what kind of dialogue you want from clues in your prompts. It is emulation, not reality.