When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent Résumé Screening

Jian Gao and Hang Jiang replace one-call resume screening with a two-agent exchange between employer-side and candidate-side agents, and measure both who advances and whether that outcome recurs.
Ask this paper
Setup. 600 constructed resume-job pairs are screened by GPT-5.5 and Claude Opus 4.7, comparing a static single judgment against an exchange in which both sides present evidence and update before deciding.
More applications advance. Advance rates rise from 33.3% to 39.3% for GPT-5.5 and 34.0% to 35.5% for Opus 4.7; on the 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%.
Decisions change in both directions. Two-agent screening also rejects applications the one-call procedure advanced, and no one-call threshold recovers the applications that two-agent screening consistently selects.
Recurrence is weaker for the new selections. Re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a one-call follow-up shows no comparable decline.
Significance. The screening procedure, not only the model, determines who reaches human review and how stable that access is across reruns.
Abstract
Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet résumé screening, the first gate, is commonly automated as a static, one-call judgment over a résumé-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed résumé-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.