More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Ziyang Xu, Tieyong Zeng and colleagues at CUHK and the Chinese Academy of Sciences test whether populations of automatically generated LLM harnesses add real specialization, or whether their extra coverage is what rerunning a single program would give anyway.
Ask this paper
Control. On 386 MATH-500 tasks, eight generated harnesses plus a baseline are compared with nine byte-identical copies of the baseline, with three executions per member.
Identical programs already add coverage. The nine identical copies yield 2.16 points of repeat-averaged oracle headroom, so oracle coverage gains alone do not prove that different programs are specialized.
Repeatable patterns are mostly weaknesses. Generated programs lose to the baseline consistently across all three repeats on 100 tasks, while consistent wins occur on only one task, and that win depends on answer extraction.
Selection gains nothing. A selector frozen on development data gains 0.00 points, and both populations reach 98.70% oracle coverage at 27 harness executions.
Proposed standard. Claims of harness diversity should show advantages that persist across executions, support a usable selection decision, and beat extra runs of a fixed program at the same inference budget. Under review at ICLR 2027.
Abstract
Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repeatable score patterns, but these chiefly reveal persistent weaknesses: losses relative to the baseline persist across all three repeats on 100 tasks, while persistent wins occur on only one task and are sensitive to answer extraction. The frozen selector gains 0.00 percentage points, and both populations reach 98.70% oracle coverage at 27 harness executions. Stable complementarity remains unresolved at three repeats. Supporting BIRD traces locate failures in mechanism implementation, activation, and output validity. Together, these findings establish why coverage and repeatability alone cannot justify claims of useful specialization. They motivate an evaluation standard for harness diversity: task advantages should persist across executions, guide usable decisions, and improve on additional fixed-program executions under matched inference budgets.