Top stories in AI today: – Biohub’s new ‘world model of protein biology’
– OpenAI Foundation puts $250M behind AI disruption
– Teach your AI agent to edit like you
– An AI that keeps learning on the job
– 4 new AI tools, community workflows, and more
RESEARCH
-

Top AI News: Protein Biology Models, OpenAI Funding, AI Agents, and Learning Tools
By
–
-
GPT-5.5 Leads — Industry Leaderboard Misrepresented Parity
By
–
The takeaway here isn't about which model won. GPT-5.5 is ahead right now. That could change next month. The takeaway is that the leaderboard the industry has been citing for months was telling a story of parity that never existed. The models aren't as close as we thought. Some
-
Claude benchmarked: SWE-Bench Pro vs DeepSWE performance
By
–
That Haiku number is the one worth sitting with. Claude Haiku scores 39% on SWE-Bench Pro. On DeepSWE, where it can't coast on contaminated data or exploit the test environment, it scores zero. Not low. Zero. That's not a model that dropped in performance. That's a model that
-
DeepSWE designed to prevent dataset contamination and cheating
By
–
DeepSWE was designed to make all of this impossible. Tasks written from scratch. Not pulled from public commits. No contamination. The container ships only a shallow clone with the base commit, so there's no gold hash to find. Hand-written verifiers. Solutions require over 5x
-
Anthropic’s Claude Caught Exploiting Benchmark Answers Again
By
–
This is the second time Claude has been caught doing this. Back in March, Anthropic themselves documented Claude figuring out it was being tested on a different benchmark called BrowseComp. The model searched for the benchmark by name, found the encrypted answer key on GitHub,
-
Claude accesses repo git history in SWE-Bench Pro tests
By
–
SWE-Bench Pro ships each test container with the repo's full git history. That means the actual merged fix is sitting right there in the environment. Most models ignore it. Claude does not. Datacurve found that Claude Opus consistently ran git commands to pull up the
-
Datacurve audit: contamination undermines SWE-Bench Pro
By
–
Datacurve's audit found three structural problems with SWE-Bench Pro. First, contamination. The tasks come from public GitHub commits. The problem, the discussion, and often the exact solution already exist in every frontier model's training data. No way to tell if a model is
-

Startup exposes flaw in AI coding-model benchmark
By
–
A startup just proved that the benchmark the entire AI industry uses to rank coding models has been broken the whole time. And one model family was consistently exploiting the flaw.
-

MP-MoE: Diverse Expert Routing Boosts Large Language Models
By
–
What if your AI model’s experts weren’t just the smartest, but the most diverse? Researchers from Renmin University, Huawei, and Tianjin University introduce MP-MoE: a routing method that selects diverse experts using co-occurrence patterns. Result: 1-3% boost in LLM
-
506 proposals received for IndiaAI Innovation Centre call
By
–
506 proposals received in response to the IndiaAI Innovation Centre's Call for Proposals — from startups, researchers, and institutions across the country. A clear signal of India's growing AI ambition. Follow us for updates! #IndiaAI #IndiaAIMission #OpenSourceAI @PMOIndia