Not sure they're doing the exact same eval, but GPT-4o reports 90.2% vs. 84.9% for Claude 3 Opus on HumanEval (
https://
openai.com/index/hello-gp
t-4o/
…). I don't expect Codestral 22B to be on par, but it can do FIM
GPT-4o vs Claude Opus: Code Generation Performance Comparison
By
–