We think this is an overactive abuse-detection system, digging in. Also working on making the terms for -p crystal clear.
@bcherny
-
Token efficiency improvements enable extended AI sessions
By
–
We've been landing token-efficiency improvements so you can do more in your sessions. If you have a session where usage felt particularly high, can you run /bug and post the feedback id? Feel free to tag me directly
-
Report AI Model Changes and Behavioral Issues
By
–
Hmm there were no behavioral/model changes recently. If you see something weird (eg. the model refusing to do something) mind running /bug and posting the feedback id? Feel free to tag me
-
Overactive AI abuse classifier requires policy clarification
By
–
This is not intentional, likely an overactive abuse classifier. Looking, and working on clarifying the policy going forward.
-
Token Counting API Support Across Bedrock Vertex Azure Anthropic
By
–
yep we do this for Bedrock, Vertex, and Azure, since they don’t have a token-counting API available yet. When using Anthropic API we use the token-counting endpoint directly
-
Claude API Usage Policy and Overages Now Clearly Defined
By
–
Yep API is fully supported, as well as overages for Claude logins. This was always the case in our terms and docs, but since it was recommended in some other products’ 3p docs, a lot of folks didn’t realize it’s not allowed. Hoping the new way is less footgunny for people.
-
Scaling High Throughput AI Inference Infrastructure Challenges
By
–
Sometimes I take for granted how quickly we can ship great product, vs how hard it is to tune a super high throughput inference + api stack. The scale makes the latter really hard. we’re working around the clock to make it better.
-
Default Settings and Token Usage in AI Systems
By
–
Everyone gets the same default, and it’s sticky when you change it. The only setting that isn’t sticky across sessions is effort=max, because it can use a lot of tokens
-
Claude Code Effort Levels Impact Model Performance Differently
By
–
This is false. We serve exactly the same models to all users. What the person in the post might be experiencing is a lower effort level vs. what the enterprise set. Claude Code users can change this anytime by running /effort. low effort = less tokens and lower intelligence,
-
Subscription optimization for AI usage patterns at scale
By
–
It's not about tokens, it's about our subscriptions being optimized for specific usage patterns. Lots of tradeoffs in building for such large scale, and one of them is optimizing systems for certain use cases and not others