OpenAI has this interesting benchmark of OpenAI's real engineering bottlenecks, where the scores have not moved since launch over a year ago. Some earlier models did even better than 5.5. I wonder what's going on here.
RESEARCH
-

Opening AI Safety Evaluations and Datasets
By
–
AI safety can't happen behind closed doors! So cool to see that the @AISecurityInst is releasing its evals, datasets, and models in the open on @huggingface, so researchers everywhere can scrutinize, reproduce, and build on them: http://huggingface.co/ai-safety-inst
-
Raw compute vs EFC: key distinction for agent learning updates
By
–
The key distinction: raw compute measures activity. EFC measures useful closed-loop learning inside the trace. That difference matters enormously for agents, because two runs with the same token count and tool calls can differ completely in whether the agent actually updates
-

New Preprint on Scaling Laws for Agent Harnesses
By
–
Agents do not scale because they spend more compute. They scale because they turn interaction into usable feedback. A sharp new preprint by Xuanliang Zhang, Dingzirui Wang, Keyan Xu, Qingfu Zhu, and Wanxiang Che introduces: Scaling Laws for Agent Harnesses via Effective
-
Open weights models more fragile than benchmarks suggest, says Mollick
By
–
I think Epoch does a great job benchmarking, but I continue to believe that open weights models are much more fragile, especially out-of-distribution, than their benchmarks indicate. Vibe-wise, I don’t think they were only 3 months behind last year or only 4 months behind today.
-

Nemotron-Labs-Diffusion: tri-mode model boosts accuracy and throughput
By
–
Can one model switch between three decoding modes to crush both accuracy and throughput? NVIDIA researchers (with Georgia Tech, HKU, and MIT) introduce Nemotron-Labs-Diffusion — a tri-mode language model that unifies autoregressive (AR), diffusion, and self-speculation decoding
-

Microsoft transforms SKILL.md into trainable object with SkillOpt
By
–

Microsoft just turned SKILL .md into a trainable object! SkillOpt is a text-space optimizer for agent skills. Instead of hand-writing or one-shot generating your SKILL .md, SkillOpt treats the skill document as the trainable external state of a frozen agent and optimizes it
-
Confirmed again: LLMs cannot handle the truth, says Marcus
By
–
Why LLMs rarely payoff—and what I have been saying literally for 7 years—confirmed yet again: LLMs can’t handle the truth. (Nor apparently can my critics, who keep saying I am “always wrong”, when I have been saying keeps being confirmed, over and over again.)
-
Seedance 2.0 still unbeaten in text-to-video since February
By
–
I still find it crazy that no lab has surpassed Seedance 2.0 in text-to-video, even though Seedance 2.0 was released back in February.
-

Perplexity enhances Daily Digest with customizable sources and connectors
By
–
Perplexity keeps working on the Daily Digest feature, allowing users to precisely customise from where and which data needs to be pulled from. Memory, web sources, custom instructions and many connectors will be available.
