AutoMix
Free while signed in. Answers cite the passages they came from.

AutoMix routes queries between LLMs of different sizes based on smaller-model confidence, saving cost without sacrificing quality.
Confidence-based routing: A small model answers first; a confidence signal determines whether to accept its answer or escalate to a larger model.
Cascading thresholds: Uses multiple confidence thresholds to route queries through a cascade of increasingly capable (and expensive) models.
Cost-quality Pareto: Achieves Pareto improvements over single-model baselines, delivering equivalent quality at substantially lower inference cost.
Production relevance: The pattern maps cleanly onto practical LLM deployment where most queries can be handled by cheap models but a tail of hard queries need the frontier model.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack