Announcing improved answer-correctness judge in Agent Evaluation, offering significant accuracy gains, especially on customer-representative use cases. Developed by @DbrxMosaicAI
's research + engineering teams. Now automatically available to all users.
Improved Answer-Correctness Judge for Agent Evaluation
By
–