🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Evaluation · Reinforcement Learning · Agents

F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows

First page
F$^{2}$DR: A Fine-Grained Full-Pipeline Reward Framework for DeepSearch Workflows
The curator’s take

Bojian Xiong and a 14-author team score a DeepSearch run across its whole pipeline rather than only its final answer, and release a benchmark for reward models in that setting.

Ask this paper

Key points
01

What existing reward models miss. They are built for static single-turn tasks, so they cannot grade the closed loop of planning and reflection, retrieval, and answer generation that DeepSearch actually runs.

02

Three scoring dimensions. F2DR evaluates Content, Trajectory and Answer separately, which yields process-level assessment instead of one outcome score.

03

DeepSearch RM-Bench. A dedicated benchmark for reward models in DeepSearch scenarios, which the experiments show discriminates well across existing open-source reward models.

04

Results. F2DR reaches substantially higher evaluation consistency than self-evaluation baselines, where the model grades its own run.

05

Significance. Agentic search systems are increasingly trained against reward models, and this is a direct measurement of whether those reward models can see the parts of the run that matter.

Abstract

With the widespread industrial deployment of Large Language Models (LLMs), DeepSearch has emerged as the dominant paradigm for resolving complex user queries. It typically operates through an iterative closed-loop workflow consisting of planning and reflection, information retrieval, and answer generation. However, existing reward models (RMs) and evaluation benchmarks are primarily designed for static single-turn tasks, failing to capture the full-pipeline complexity of DeepSearch workflows. To address this limitation, we propose F2DR, a fine-grained full-pipeline DeepSearch reward framework. F2DR evaluates DeepSearch workflows across three dimensions: Content, Trajectory, and Answer, enabling comprehensive process-level assessment. We further construct DeepSearch RM-Bench, a dedicated benchmark for evaluating RMs in DeepSearch scenarios. Extensive experiments demonstrate that F2DR achieves significantly higher evaluation consistency than self-evaluation-based baselines, while DeepSearch RM-Bench exhibits strong discriminative capability across existing open-source RMs. We will publicly release the complete DeepSearch RM-Bench dataset soon.

Every Monday
Get next week’s papers.
Subscribe on Substack