Man, what's going on with this benchmark, it is literally getting worse with each release. "OpenAI-Proof Q&A evaluates AI models on 20 internal research and engineering bottlenecks encountered at OpenAI, each representing at least a one-day delay to a major project and in some
RESEARCH
-

CGM AI Superior to Self-Monitoring for Type 2 Diabetes Control
By
–
A randomized trial of continuous glucose monitoring (CGM) vs self-monitoring for Type 2 diabetes on basal insulin and drug therapies shows superiority of CGM for gluocse regulation https://
thelancet.com/journals/landi
a/article/PIIS2213-8587(26)00076-8/fulltext
… @TheLancetEndo -
DeepLearning.AI optimizes cost accuracy latency balance
By
–
Are you coming to @DeepLearningAI AI's AI Dev 26 x SF next week?
— AI21 Labs (@AI21Labs) 23 avril 2026
Our Chief Product & Strategy Officer, Or Dagan, will be giving a talk diving into what our product and R&D teams are doing to automate finding the optimal balance between cost, accuracy, and latency for every… pic.twitter.com/GrWaVO58bEAre you coming to @DeepLearningAI AI's AI Dev 26 x SF next week? Our Chief Product & Strategy Officer, Or Dagan, will be giving a talk diving into what our product and R&D teams are doing to automate finding the optimal balance between cost, accuracy, and latency for every
-

GPT 5.5 vs Mythos: Benchmark Performance Comparison Analysis
By
–
A false narrative is being shared that GPT 5.5 ties with Mythos on several benchmarks as if that made them equivalent. What's being overlooked is that GPT 5.4 was already on par with Mythos in those benchmarks, except for Terminal Bench 2.0…
-
Testing AI Model 5.5 Beyond Benchmark Limitations
By
–
yep. I think benchmarks only tell half of the truth. Going to test 5.5 now
-

GPT-5.5 Pro Reaches Claude Mythos Level Performance
By
–
From an eval perspective, GPT-5.5 pro is Claude Mythos level but for public use.
-
AI Model Selects Correct Method from Research Paper Appendix
By
–
45 → 65 is something! crazy that it picked the right method from a paper appendix. I know many PhDs that dont even spend the time going over the appendices haha
-

OpenAI Releases GPT-5.5 Most Capable Model Yet
By
–
GPT-5.5 evals are crazy – and what we have hoped for.
— Chubby♨️ (@kimmonismus) 23 avril 2026
Whats new? the tl;dr
OpenAI just dropped GPT-5.5, their "most capable model yet", which matches GPT-5.4's latency while delivering significantly higher intelligence across coding, knowledge work, and scientific research.… https://t.co/fDII5T1KeV pic.twitter.com/qdhXhPiWSfGPT-5.5 evals are crazy – and what we have hoped for. Whats new? the tl;dr OpenAI just dropped GPT-5.5, their "most capable model yet", which matches GPT-5.4's latency while delivering significantly higher intelligence across coding, knowledge work, and scientific research.
-

CORES Symposium 2026: AI and Scientific Reproducibility
By
–
As AI in science gains momentum, this year's CORES Symposium brings together researchers and data scientists on April 29 to tackle AI's role in scientific reproducibility. The event is sold out, but you can still join the waitlist here: https://
datascience.stanford.edu/events/open-sc
ience-center/cores-annual-symposium-2026-ai-and-scientific-reproducibility
… -

Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
By
–
Fine-Tuning Diffusion Models via Intermediate Distribution Shaping paper: https://
huggingface.co/papers/2510.02
692
…
