AI Dynamics

Global AI News Aggregator

About

Kimi outperforms GPT-5 on key benchmarks like Humanity’s Last Exam

the benchmarks aren't close. they're embarrassing. > humanity's last exam: kimi 44.9%, beats closed models
> browsecomp (agentic web search): kimi 60.2% vs gpt-5's 54.9%
> gpqa diamond: 85.7% vs gpt-5's 84.5%
> swe-bench verified: 71.3% (coding tasks)
> artificial analysis

→ View original post on X — @godofprompt