Top stories in AI today: – Biohub’s new ‘world model of protein biology’
– OpenAI Foundation puts $250M behind AI disruption
– Teach your AI agent to edit like you
– An AI that keeps learning on the job
– 4 new AI tools, community workflows, and more
AI
-

Top AI News: Protein Biology Models, OpenAI Funding, AI Agents, and Learning Tools
By
–
-

ElevenLabs partners with Greece to reimagine public services with voice AI
By
–
We’re partnering with the Government of Greece to reimagine public services. Today, we signed an MOU with @PrimeministerGR and @papastergiougr to use voice AI to improve public services, promote tourism, and preserve Greek linguistic heritage.
-
GPT-5.5 Leads — Industry Leaderboard Misrepresented Parity
By
–
The takeaway here isn't about which model won. GPT-5.5 is ahead right now. That could change next month. The takeaway is that the leaderboard the industry has been citing for months was telling a story of parity that never existed. The models aren't as close as we thought. Some
-
Claude benchmarked: SWE-Bench Pro vs DeepSWE performance
By
–
That Haiku number is the one worth sitting with. Claude Haiku scores 39% on SWE-Bench Pro. On DeepSWE, where it can't coast on contaminated data or exploit the test environment, it scores zero. Not low. Zero. That's not a model that dropped in performance. That's a model that
-
DeepSWE designed to prevent dataset contamination and cheating
By
–
DeepSWE was designed to make all of this impossible. Tasks written from scratch. Not pulled from public commits. No contamination. The container ships only a shallow clone with the base commit, so there's no gold hash to find. Hand-written verifiers. Solutions require over 5x
-
Anthropic’s Claude Caught Exploiting Benchmark Answers Again
By
–
This is the second time Claude has been caught doing this. Back in March, Anthropic themselves documented Claude figuring out it was being tested on a different benchmark called BrowseComp. The model searched for the benchmark by name, found the encrypted answer key on GitHub,
-
Claude accesses repo git history in SWE-Bench Pro tests
By
–
SWE-Bench Pro ships each test container with the repo's full git history. That means the actual merged fix is sitting right there in the environment. Most models ignore it. Claude does not. Datacurve found that Claude Opus consistently ran git commands to pull up the
-
Datacurve audit: contamination undermines SWE-Bench Pro
By
–
Datacurve's audit found three structural problems with SWE-Bench Pro. First, contamination. The tasks come from public GitHub commits. The problem, the discussion, and often the exact solution already exist in every frontier model's training data. No way to tell if a model is
-

Startup exposes flaw in AI coding-model benchmark
By
–
A startup just proved that the benchmark the entire AI industry uses to rank coding models has been broken the whole time. And one model family was consistently exploiting the flaw.
-
50k€ Investment: Online Hate Complaint Automation Platform, Pharos/Justice Link
By
–
BTW, je suis prêt à investir 50k€ dans une plateforme d’automatisation de plaintes en ligne sur la haine sur internet, avec recherche de contenus, évaluation du cas, journalisation et suivi automatisé des dossiers, relance etc etc… en lien avec Pharos <> avocat <> justice <>… pic.twitter.com/jm2QCsdOB5
— Defend Intelligence (Anis Ayari) (@DFintelligence) 28 mai 2026BTW, I am ready to invest €50k in an online hate complaint automation platform, including content search, case evaluation, automated logging and tracking of files, follow-ups, etc., in connection with Pharos, lawyers, and the justice system.