AI Dynamics

Global AI News Aggregator

About

Comparative Performance Analysis of LLMs Across Task Complexity

How do models perform? Human: 91%+
GPT-5: ~54%
GPT-4.1: ~36%
GPT-4o: ~52%
DeepSeek-R1: ~55% And here’s the kicker: 3-app tasks cause up to 3x degradation vs single-app tasks across all models.

→ View original post on X — @godofprompt