.
@Harvey
’s LAB benchmark approaches verification like a human would. Every task in a dataset has criteria for the task to pass. Legal agents can have 50+, with each one having its own judge call. It’s easy to audit, but can be expensive at scale. LangChain Labs teamed up with
Harvey’s LAB benchmark uses human-like verification with per-task criteria
By
–
