AI Dynamics

Global AI News Aggregator

About

Evaluation of LLM Performance Gains in Reasoning and Agentic Workflows

Math got much better for sure. Coding improved too, but less good at SVGs. They are also testing a lot of agentic use cases – working with docs, health data, etc.

→ View original post on X — @testingcatalog