6). An Open-source LM Specialized in Evaluating Other LMs – open-source Prometheus 2 (7B & 8x7B), state-of-the-art open evaluator LLMs that closely mirror human and GPT-4 judgments; they support both direct assessments and pair-wise ranking formats grouped with user-defined
Prometheus 2: Open-Source LM for Evaluating Language Models
By
–