AutoMix
First page

Paper summary
AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.
Ask this paper
01
Confidence-based routing: A small model answers first; a confidence signal determines whether to accept its answer or escalate to a larger model.
02
Cascading thresholds: Uses multiple confidence thresholds to route queries through a cascade of increasingly capable (and expensive) models.
03
Cost-quality Pareto: Achieves Pareto improvements over single-model baselines, delivering equivalent quality at substantially lower inference cost.
04
Production relevance: The pattern maps cleanly onto practical LLM deployment where most queries can be handled by cheap models but a tail of hard queries need the frontier model.