More Agents Is All You Need
Free while signed in. Answers cite the passages they came from.

The paper shows that simply running more independent LLM agents and voting produces reliable scaling gains across tasks, without any method changes.
Sampling-and-voting: For a given task, run N independent LLM agents on the same query, then majority-vote over their answers - a minimalist ensemble.
Scales with agent count: Performance improves monotonically with more agents across reasoning, coding, and QA benchmarks, with larger gains on harder problems.
Orthogonal to other tricks: The gains stack on top of existing improvements like prompt engineering, CoT, and RAG, making ensembling a free-standing lever.
Implication: Raw parallel ensembling is surprisingly strong compared to architecturally complex multi-agent systems and should be a baseline in any comparative study.
Get next week’s papers.
The same picks and the same summaries, in your inbox. Free, and 176 issues deep.
Subscribe on Substack