Oh look, AI folks have discovered stacking!
@tunguz
-

NVIDIA GB300 outperforms H200 by 20x in agents per megawatt
By
–
Among other findings the benchmarks shows that NVIDIA's GB300 outperforms H200 by a whopping factor of 20 in terms of running concurrent agents per megawatt. 3/4
-
AA methodology and Nvidia technical blog on agentic coding performance
By
–
AA methodology: https://
artificialanalysis.ai/methodology/ag
entperf
…
Nvidia technical blog: https://
developer.nvidia.com/blog/nvidia-ac
hieves-leading-agentic-coding-performance-on-first-agentic-ai-benchmark/
… 4/4 -

Artificial Analysis launches AgentPerf, first agentic AI benchmark
By
–
Artificial Analysis just announced AgentPerf, the industry’s first agentic AI benchmark. The benchmark uses several real-world agentic use trajectories and employs OpenCode agentic harness using three top open-source models with reasoning enabled 1/4
-
DeepSeek, GLM, Kimi solve real code issues with reasoning and tools
By
–
(DeepSeek V3.2, GLM 4.7, and Kimi K2.5) prompted to resolve issues in real public code repositories. All trajectories include interleaved reasoning and tool calls. 2/4
-
US faces prospect of sabotaging its AI export dominance
By
–
AI was ONE huge export product that the US absolutely dominated over the past few years. Now we are facing the prospect of completely and utterly sabotaging it.
-
Fully solving coding shows it wasn’t the hard part
By
–
Once we *fully* solve coding, we’ll realize that coding wasn’t the hard part.
-

Claude model unpredictability compared to a box of chocolates
By
–
Claude is like a box of chocolates. You never know which model are you going to get.
-

GPT 5.6 likely drops June 23, potentially harming Anthropic
By
–
I’m now 99.99% certain that GPT 5.6 will drop, in Codex and everywhere else, on June 23rd. This could end up being Anthropic’s biggest self own.
-
Timeline for open source model reaching Mythos level
By
–
How soon do you think we'll have an open source model at the level of Mythos?
