🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

An Open-source LM Specialized in Evaluating Other LMs

First page
An Open-source LM Specialized in Evaluating Other LMs
Paper summary

open-source Prometheus 2 (7B & 8x7B), state-of-the-art open evaluator LLMs that closely mirror human and GPT-4 judgments; they support both direct assessments and pair-wise ranking formats grouped with user-defined evaluation criteria; according to the experimental results, this open-source model seems to be the strongest among all open-evaluator LLMs; the key seems to be in merging evaluator LMs trained on either direct assessment or pairwise ranking formats.

Ask this paper

Every Monday
Get next week’s papers.
Subscribe on Substack