Evaluations with No Labels
Free while signed in. Answers cite the passages they came from.

Self-supervised evaluation of LLMs via sensitivity/invariance to input transformations.
Label-free evaluation: Evaluates LLMs without requiring ground-truth labels, using consistency under input perturbations as the signal.
Transformation-based probes: Measures sensitivity or invariance to paraphrasing, irrelevant-context addition, and other transformations that shouldn't change correct answers.
Live deployment monitoring: Useful for monitoring LLM behavior on datasets streamed during production deployment, catching drift without manual labeling.
Deployment infrastructure: An early contribution to the continuous evaluation tooling that would become standard for 2024 LLM production systems.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack