🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
Evaluation · Code · Agents

Harness-IF

Free while signed in. Answers cite the passages they came from.

First page
Harness-IF
The curator’s take

When a coding agent obeys your rule, it may simply have been going to do that anyway. Existing instruction-following benchmarks cannot tell compliance from coincidence because they concentrate rules in the user turn, while coding-agent benchmarks only score final task success.

Key points
01

Rules are the unit, not tasks: A 642-rule library places 302 rules across the five configurable surfaces a deployed agent actually reads, spread over 60 realistic multi-turn coding items, with 256 rules receiving execution-grounded verdicts one at a time.

02

A metric that strips out luck: Against-Prior Accuracy scores only rules labeled as opposing unprompted defaults, established by re-running tasks with the rule withheld across nine probe builds. Across 12 frontier models, raw accuracy spans 72.1 to 85.9% and AP-Acc spans 66.1 to 78.6%.

03

The inflation is model-specific: Every model is worse on against-prior rules, by 3.6 to 7.4 points with a mean of 5.81, and the inflation varies twofold across the cohort, so aggregate scores are not comparable between builds without the correction.

04

Why it matters: A counterbalanced conflict pilot on nine separate builds finds that precedence does not follow prompt depth. System prompts, project files, and user instructions all outrank tool and skill descriptions, which is worth knowing before you decide where to put a rule you actually need followed.

Every Monday
Get next week’s papers.

The same picks and the same summaries, in your inbox. Free, and 176 issues deep.

Subscribe on Substack