Evaluations with No Labels
First page

Paper summary
Self-supervised evaluation of LLMs via sensitivity/invariance to input transformations.
Ask this paper
01
Label-free evaluation: Evaluates LLMs without requiring ground-truth labels, using consistency under input perturbations as the signal.
02
Transformation-based probes: Measures sensitivity or invariance to paraphrasing, irrelevant-context addition, and other transformations that shouldn't change correct answers.
03
Live deployment monitoring: Useful for monitoring LLM behavior on datasets streamed during production deployment, catching drift without manual labeling.
04
Deployment infrastructure: An early contribution to the continuous evaluation tooling that would become standard for 2024 LLM production systems.