The effort default change from high to medium that Anthropic confirmed is real and measurable. A 15-point accuracy drop in days on a specific benchmark needs a citation.
LLMS
-
Llama 4 Launch Drives Meta AI Downloads and Chart Movement
By
–
The Llama 4 launch generated enough press coverage to drive curiosity downloads from people who had never tried Meta AI before. That's enough to move the charts temporarily.
-
Anthropic Changes Claude Default Effort Level to Medium
By
–
Anthropic confirmed the effort default changed from high to medium. That's not a nerf for its own sake, it's a cost and compute management decision with real tradeoffs for users.
-
Old School VS Coder Supercharged with Claude AI
By
–
I’m old school VS coder supercharged with Claudio.
-

LLMs Fairness and Consistency in AI Model Evaluation
By
–
Are LLMs truly fair and consistent when judging other AI models? A collaborative team from Peking University, NUS, Institute of Science Tokyo, Nanjing University, Carnegie Mellon, Westlake, and Southeast University has the answer! They introduce TrustJudge, a probabilistic
-

Context length benchmarks up to 180k tokens
By
–
Ran some benchmarks on different context lengths all the way to 180k Findings below
-

Ways to Train Large Language Models Effectively
By
–
Ways to Train an #LLM
by @goyalshaliniuk #GenAI #ArtificialIntelligence #MachineLearning #ML -

Exploring dense new format with tests in BF16, FP8, NVFP4 and TP scaling
By
–

it's very dense, a new format that I am playing with still need to test it in concurrency + try it in nvfp4 hopefully will be able to compare perplexity in bf16 / fp8 / nvfp4 as well as TP performance jumps from 2 -> 4 -> 8 nodes across all 3 formats most important thing
-

Anthropic Developing Coordinator Mode for Multi-Agent Delegation
By
–
Anthropic is working on a new Coordinator Mode, along with the ability to create custom agents directly from the desktop app. > "Delegate work across parallel agents." Will Coordinator Mode be similar to an upcoming Magic TODO list on Codex? h/t @M1Astra