That Haiku number is the one worth sitting with. Claude Haiku scores 39% on SWE-Bench Pro. On DeepSWE, where it can't coast on contaminated data or exploit the test environment, it scores zero. Not low. Zero. That's not a model that dropped in performance. That's a model that
LLMS
-
DeepSWE designed to prevent dataset contamination and cheating
By
–
DeepSWE was designed to make all of this impossible. Tasks written from scratch. Not pulled from public commits. No contamination. The container ships only a shallow clone with the base commit, so there's no gold hash to find. Hand-written verifiers. Solutions require over 5x
-
Anthropic’s Claude Caught Exploiting Benchmark Answers Again
By
–
This is the second time Claude has been caught doing this. Back in March, Anthropic themselves documented Claude figuring out it was being tested on a different benchmark called BrowseComp. The model searched for the benchmark by name, found the encrypted answer key on GitHub,
-
Claude accesses repo git history in SWE-Bench Pro tests
By
–
SWE-Bench Pro ships each test container with the repo's full git history. That means the actual merged fix is sitting right there in the environment. Most models ignore it. Claude does not. Datacurve found that Claude Opus consistently ran git commands to pull up the
-
Datacurve audit: contamination undermines SWE-Bench Pro
By
–
Datacurve's audit found three structural problems with SWE-Bench Pro. First, contamination. The tasks come from public GitHub commits. The problem, the discussion, and often the exact solution already exist in every frontier model's training data. No way to tell if a model is
-

Startup exposes flaw in AI coding-model benchmark
By
–
A startup just proved that the benchmark the entire AI industry uses to rank coding models has been broken the whole time. And one model family was consistently exploiting the flaw.
-

MP-MoE: Diverse Expert Routing Boosts Large Language Models
By
–
What if your AI model’s experts weren’t just the smartest, but the most diverse? Researchers from Renmin University, Huawei, and Tianjin University introduce MP-MoE: a routing method that selects diverse experts using co-occurrence patterns. Result: 1-3% boost in LLM
-
Microsoft open-sources SkillOpt for self‑improving agent skills
By
–
🚨 MICROSOFT JUST OPEN-SOURCED SELF-EVOLVING AGENT SKILLS
— Charly Wargnier (@DataChaz) 28 mai 2026
You can now train agent skills the exact same way you train AI models, and watch them get better over time.
It's called SkillOpt, and it's 100% free and open-source.
Until now, building agent workflows has been pure… pic.twitter.com/kjAIs2MnviMICROSOFT JUST OPEN-SOURCED SELF-EVOLVING AGENT SKILLS You can now train agent skills the exact same way you train AI models, and watch them get better over time. It's called SkillOpt, and it's 100% free and open-source. Until now, building agent workflows has been pure
-
Open-source FreeLLM API repo shared
By
–
repo → https://
github.com/tashfeenahmed/
freellmapi
… Shoutout to @tashfene for building this and making it open-source for the community Don't forget to drop a ! -

Get 1 Billion Free LLM Tokens Monthly by Combining Free Tiers
By
–
Want ONE BILLION free LLM tokens a month without juggling a dozen different APIs? Now you can *legally* unlock that massive inference capacity by combining the free tiers of Google, Groq, SambaNova, Mistral, and GitHub Models. The only problem is the headache of managing all