
Absolutely incredible: GLM-5.2 (max) places 3rd overall on GDPval-AA, a real-world agentic work benchmark, even ahead of GPT-5.5 (xhigh). Oh and by the way: it seems open source is no longer 7 months behind. GDPval-AA, a benchmark built around tasks.

