Blending Is All You Need

Small chat models (6B/13B) blended together can rival ChatGPT-class systems, without any new training.
Ask this paper
Blending ≠ ensembling: The system samples responses from a pool of different small models per turn, letting the conversational mix drive quality and diversity rather than any single dominant model.
ChatGPT-comparable engagement: Evaluations on a production chat platform show that a Blended system of 6B + 13B models achieves engagement comparable to ChatGPT, despite having far fewer parameters and compute.
Diversity win: The mix of different small models produces more varied responses than a single larger model, likely driving a portion of the engagement gains.
Practical recipe: Suggests that teams without frontier-scale budgets can compose existing open models to deliver competitive chat experiences, repositioning "bigger is better" thinking for production chat.