AI Dynamics

Global AI News Aggregator

About

Mistral Large 4 benchmark scores across coding, finance, legal tasks

Mistral Large 4 scores 62% on DeepSWE, outperforming GLM-5.3, according to VentureBeat. Additionally, it scores 67% on Finch (Financial tasks, SOTA open-weight). Le Chonk also scores 15% on Harvey’s Legal Agent Benchmark (Legal tasks, SOTA open-weight). We need a tech report

→ View original post on X — @testingcatalog