Whatever people build around the model today becomes training data for the next one, that loop just keeps tightening to, TBH, a worrying point haha. Quite excited/scared to see where this "slopacopalypse" will lead us.
RESEARCH
-
Semantic Correctness Over Exact Match for Agent Evaluation
By
–
Semantic correctness over exact match is the right call for agent work, exact match penalizes paraphrases the downstream agent doesn't even care about.
-
Medical imaging AI proves effective but LLMs lack real world evidence
By
–
Superhuman interpretation AI for medical images, such as mammography and endoscopy, has been proven to improve diagnostic accuracy in multiple randomized trials but mostly not implemented. But LLMs for clinical decision support have little real world medicine proof, but are
-
Opus 4.7 Performance Issues: Major Regression Compared to GPT-4o
By
–
i'm a few days late to realizing this but: wow, opus 4.7 is god awful like so, so bad it's making mistakes on things i'd expect gpt-4o to handle cleanly there's got to be some explanation, right?
-

GPT-5.5 achieves SOTA performance on long-running tasks
By
–
SOTA perf from GPT-5.5 for long-running tasks:
-

Open Benchmark Grants Initiative Funds Frontier AI Research
By
–
We're proud to have supported this work through Open Benchmark Grants, our initiative funding the next wave of frontier benchmarks!
-
GPT-5.5 Released: Major Performance Improvements Available
By
–
gpt-5.5 is a big step up in performance, give it a try:
-
Are LLMs Really More Important Than Fire or Electricity?
By
–
Are LLMs really more important than fire or electricity? “Honestly, a ton of what we’ve developed in my lifetime amounts to scaling up the delivery of information and entertainment and the frictionlessness of certain financial transactions. These are real improvements! … But
-

Token Importance in On-Policy Distillation with Selective Training
By
–
"TIP: Token Importance in On-Policy Distillation" This paper introduces selective token training for on-policy distillation, relying on student entropy to find high-signal tokens. A key point is that entropy misses confident mistakes, so they add teacher-student divergence to