🚀NEW LABGetting Started with Claude AgentsStart lab
← All papers  /  Oct 9, 2026
Agents

Humanize: Judgement Engineering for Agentic Coding

First page
Humanize: Judgement Engineering for Agentic Coding
The curator’s take

Sihao Liu, Ligeng Zhu, Song Han and colleagues at NVIDIA (with UCLA, MIT and Tsinghua) describe Humanize, an open-source builder-reviewer workflow for agentic coding in which deterministic hooks, not a model, decide when work moves between planning, implementation, review and learning.

Ask this paper

Key points
01

Workflow. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from a different vendor decides completion. Hooks enforce 72 mechanical gates between the roles.

02

Joint sampling. Modeled as a Markov chain over repository states, alternating builder and reviewer samples from two models, so a defect survives only if both models miss it.

03

Applications. A 567-file gem5 build-system migration under upstream review; Kernel Design Agents placing top three in all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; 672/672 on PutnamBench and first place (251/303) on Lean-Eval; full scores reported on IOI, IMO and IPhO 2026.

04

Postmortems. In 118 public postmortems of real loops, independent review catches builder claims the evidence does not support, but stopping is the weak point: where rounds are separated by phase, two thirds happened after implementation had already been accepted.

05

Caveat. The evidence is observational deployment data (68 versions in 108 days), not a controlled comparison against other workflows.

Abstract

Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in rounds, and a reviewer agent from another vendor decides completion; deterministic hooks, not a model, route work between these roles and enforce 72 mechanical gates. Viewed as a Markov chain over repository states, alternating builder and reviewer samples jointly from two models, so a defect survives only if both miss it. We study Humanize through its deployment, 118 public postmortems of real loops, and its applications. Over 68 versions in 108 days, it gathered 1,468 GitHub stars. Applications include a 567-file gem5 build-system migration under upstream review; Kernel Design Agents, which extend the loop with a kernel knowledge base and profiling feedback and placed in the top three of all three Full-Agent tracks of the MLSys 2026 FlashInfer contest; and, through Humanize Olympiad Agents (HOA), full scores in IOI 2026, IMO 2026, IPhO 2026, and IBO 2024, 418.5/437 in IChO 2026 (gold-medal). Humanize also achieves 672/672 on PutnamBench and ranks first (251/303) on Lean-Eval's leaderboard even competiting with professional mathematicians. The postmortems show that independent review catches unsupported builder claims, but stopping remains a key weakness. In reports that separate rounds by phase, two thirds of rounds occurred after implementation was accepted. This evidence is observational, not a controlled comparison of workflows.

Every Monday
Get next week’s papers.
Subscribe on Substack