🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Aug 29 – Aug 29, 2026
Agents · Training

Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs

First page
Retrieve, Match, Escalate: Accurate and Scalable Product Linking with VLM-Distilled Cross-Encoders and Agentic VLMs
The curator’s take

Jian Wang and colleagues describe a production product-linking cascade at marketplace scale where a distilled cross-encoder auto-resolves the easy majority and an agentic VLM with web search settles only the ambiguous tail.

Ask this paper

Key points
01

Cost spans five orders of magnitude per pair: From the cheap cross-encoder to the frontier VLM. The entire design question is which fraction of traffic reaches which tier.

02

Distilled from dual-VLM consensus labels: Millions of consensus labels retire human annotation from the training set, and the cross-encoder is calibrated to auto-accept at a 98 percent precision bar validated against an operator-certified audit.

03

Open-weight agent at one seventh the cost: The self-hosted agent reaches a closed frontier VLM's precision at a four-point recall cost, 88 versus 92 percent, with no fine-tuning.

04

Escalation raises coverage from 68 to 77 percent: Sending only the hard tail to the agent is what buys end-to-end link coverage, not upgrading the whole pipeline.

05

Why it matters: This is a clean published example of difficulty-proportional compute in production, which is the pattern most agent cost discussions gesture at without numbers.

Abstract

Product linking, the entity-resolution task of mapping merchant product records to canonical catalog products, consolidates fragmented listings so downstream search, recommendation, and advertising see one clean entry per product. At marketplace scale, billions of noisy, multi-category records must be resolved against tens of millions of canonical products, where scoring every candidate with a single model is either too weak for the hard cases or too costly for the easy ones. We present a production retrieve-then-match cascade that spends computation in proportion to difficulty: retrieval surfaces plausible matches, a lightweight text cross-encoder auto-resolves the high-confidence majority, and an agentic multimodal vision-language model settles the ambiguous remainder by inspecting product images and issuing web searches for evidence that is in neither record. The cross-encoder is distilled from millions of dual-VLM-consensus labels, retiring human annotation from the training set, and is calibrated to auto-accept links at a 98% precision bar validated against a smaller operator-certified audit. The agent is a self-hosted open-weight model that reaches a closed frontier VLM's precision at a four-point recall cost (88% versus 92%) for roughly one-seventh the per-pair cost, with no fine-tuning. Per-pair cost spans nearly five orders of magnitude from the cheap cross-encoder to the frontier VLM, so escalating only the hard tail to the agent raises end-to-end link coverage from the cheap stage's 68% to 77%.

Every Monday
Get next week’s papers.
Subscribe on Substack