FrugalGPT
First page

Paper summary
Strategies to reduce LLM inference cost while improving performance.
Ask this paper
01
Three-layer strategy: Combines prompt adaptation, LLM approximation, and LLM cascading to save cost.
02
Model cascade: Routes easy queries to cheap models and escalates to expensive models only when needed.
03
Cost reduction: Shows 98% cost savings while sometimes improving accuracy over using the most expensive model always.
04
Production patterns: Influenced production LLM routing patterns and the 2024 ecosystem of LLM routers (RouteLLM, Martian).