From isolated snippets → full workflows. Proud to partner with @Stanford and @LaudeInstitute on #TerminalBench 2.0 — helping redefine agent evaluation. Thanks @Mike_A_Merrill and @alexgshaw for leading the charge and allowing our researchers the opportunity to contribute.
TerminalBench 2.0: Redefining Agent Evaluation With Stanford Partnership
By
–
