If you run Claude Code in the terminal, scroll down a bit after running /usage and you should see a complete detail of specific skills, MCP, and plugins that consume your tokens. Most of the time, it's a bad plugin that causes the
LLMS
-
Larger base model capacity improves training data memorization
By
–
Yes, because the base model has far more capacity than all the previous ones, so it's better at memorizing the training dataset.
-
Allegation that Anthropic nerfed models and Dario sabotaged codebase
By
–
Imagine how long Anthropic have had their models nerfed on purpose Dario masterclass in sabotaging your codebase while you were thinking Claude Code is just working for you lol
-

Gemini 3.5 Pro and GPT-5.6 near release; Anthropic frontier lab
By
–



It's already June 9th, and Gemini 3.5 Pro and GPT-5.6 are nearing release (Google even already announced 3.5 Pro during i/o) Rumor has it that GPT-5.6 will be released as early as next week. So far, it's safe to say that – guardrails aside – Anthropic is truly the frontier lab
-
TIL: AgentsView to Calculate Tokens for Claude Fable 5
By
–
A TIL on using http://agentsview.io to calculate token cost with Claude Fable 5 despite the fact that this model is not yet included in the AgentsView pricing database https://til.simonwillison.net/llms/agentsview-custom-model-price …
-
OpenAI criticized for losing lead and credibility to Anthropic
By
–
OpenAI is hosed. There is almost no reason to prefer them over Anthropic, they have lost their lead (despite every advantage in the world), they have made commitments far far beyond their means, their leader has lost credibility, and they don’t have a plausible plan for becoming… https://t.co/9fDsAplXB2
— Gary Marcus (@GaryMarcus) 9 juin 2026OpenAI is hosed. There is almost no reason to prefer them over Anthropic, they have lost their lead (despite every advantage in the world), they have made commitments far far beyond their means, their leader has lost credibility, and they don’t have a plausible plan for becoming
-

Gary Marcus: Justine’s answer sounds like AI slop, not AGI
By
–
justine may kid about being bullied, but the answer sure sounds more like 2026 LLM-powered-AI slop than AGI to me.
-

On-Policy Distillation geometry: fewer weight updates, preserves structure
By
–
“On the Geometry of On-Policy Distillation” OPD is not just SFT mixed with RLVR. It has its own update geometry. This paper shows that OPD updates fewer weights than SFT and preserves pretrained structure better, while staying less constrained than RLVR. The key finding is
-

FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention
By
–
“FlashMemory-DeepSeek-V4: Lightning Index Ultra-Long Context via Lookahead Sparse Attention” With how Long context LLMs are being bottlenecked by KV cache, because every old token keeps consuming GPU memory even when most of it is irrelevant, this paper turns long context into