🚀NEW LABGetting Started with Claude AgentsStart lab
Evaluation

Evaluations with No Labels

First page
Evaluations with No Labels
Paper summary

Self-supervised evaluation of LLMs via sensitivity/invariance to input transformations.

Ask this paper

Key points
01

Label-free evaluation: Evaluates LLMs without requiring ground-truth labels, using consistency under input perturbations as the signal.

02

Transformation-based probes: Measures sensitivity or invariance to paraphrasing, irrelevant-context addition, and other transformations that shouldn't change correct answers.

03

Live deployment monitoring: Useful for monitoring LLM behavior on datasets streamed during production deployment, catching drift without manual labeling.

04

Deployment infrastructure: An early contribution to the continuous evaluation tooling that would become standard for 2024 LLM production systems.

Every Monday
Get next week’s papers.
Subscribe on Substack