Traverse: Learning When to Remember, Reset, and Redirect for Long-Horizon Web Search

Jingyuan Ma, Zhifang Sui and colleagues at ByteDance and Peking University introduce Traverse, a search harness in which the agent manages its own process through Rubric, Answer and Verify states and compresses its context with a Seal Memory tool.
Ask this paper
Harness. The agent first writes criteria for a valid answer, searches under them, then verifies the result independently before deciding to stop or continue.
Seal Memory and Seal Collapse. The agent can seal parts of its context. Training this with RL caused unstable training (Seal Collapse); training only the final segment after each context-management step fixes it.
Results. The 35B model scores 72.83 on BrowseComp, above comparable open-source systems, and improves over its base model on BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation and product search.
Ablation. Agent-controlled compression outperforms automatic compaction.
Abstract
Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.