I'm tired of models being over-optimized for benchmarks but not actually being better. I think this one is one of those 'actually better' ones, especially for non-coding/reasoning tasks like writing
@thatroblennon
-
When Friends Have Hallucinations We Call It Being Wrong
By
–
and when a friend has a hallucination we call it "being wrong"
-
Using Anthropic Console’s Built-in Tool for Quick Megaprompts
By
–
for a quick megaprompt, I like the build-in tool in console (dot) anthropic (dot) com. you've got to have billing info in the system to use it, but it's free to have it write a prompt for you
-
Windsurf Usage Features and Knowledge Files Explanation Request
By
–
Yeah if you are anyone could explain in concrete specific terms how to use Windsurf in a way that's different, would love that. They seem to have some self-building knowledge files and some goodies in that respect, but I haven't used it since those upgrades
-
Tried it for a day but couldn’t tell it apart from Cursor
By
–
I played with it for a day but couldn't figure out how it was different from Cursor. It just seemed like the same thing to me
-
AI Capability Levels: From Tool Use to Autonomous Agents
By
–
I would put tool use and multi-prompt chains in here for level 1. Then maybe reasoning models at level 2. No actual agency until level 3, where I'd put the OODA-loop style autonomous agents. Then 4 and 5 are more like sci-fi AI intelligence
-

Claude Sonnet 3.7 Hidden in Amazon Bedrock Ahead of AWS Event
By
–
Sources are reporting Claude Sonnet 3.7 is hidden in Amazon Bedrock and likely to be announced during the AWS event in 2 days. (Amazon made a huge investment into Anthropic a ways back.) This is anticipated to be a reasoning + computer use model better at coding than 3.5
-
Full Self Driving Agents vs Cruise Control Assist Agents
By
–
Full Self Driving Agents vs Cruise Control Assist Agents?