🚀NEW COURSEVibe Coding AI Apps with Claude Code 🤖✨Enroll now
← All papers  /  Sep 20, 2026
Evaluation · Agents

Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions

First page
Replan, Repair, or Edit? A Unified Empirical Evaluation of Travel Agents for Itinerary Revision under Resource Disruptions
The curator’s take

Xiaofei Yuan and colleagues compare full replanning, classical plan repair and LLM-based local revision on the same disrupted-itinerary benchmark, which prior work could not do because each method defined the task differently.

Ask this paper

Key points
01

Two benchmark sets from TREK. 500 single-disruption cases spanning feasible and infeasible instances, plus 200 feasible compound-disruption cases, evaluated on effectiveness, plan stability and cost.

02

Three systems, one protocol. LLM-Z3 full replanning, IPyHOPPER hierarchical repair and an iTIMO local-revision adapter are measured under identical conditions.

03

Results. LLM-Z3 with Gemini reaches the highest compound-disruption success; IPyHOPPER nearly matches its single-disruption success while preserving substantially more of the accepted itinerary on successful repairs.

04

Cost profiles diverge sharply. IPyHOPPER uses no LLM inference at all, the LLM-Z3 adapter uses compact one-call inference, and the iTIMO adapter consumes substantially more tokens.

05

Significance. Commitment preservation and feasibility recovery pull in different directions, and the paper gives the first apples-to-apples numbers for choosing between them.

Abstract

Travel-planning agents generate itineraries that may become infeasible after acceptance because of flight cancellations, hotel unavailability, or attraction closures. Revising these itineraries involves full replanning, classical plan repair, and LLM-based travel-agent revision, whose differing task formulations and evaluation protocols hinder comparison. We conduct a systematic empirical study using two TREK-derived benchmark sets: 500 single-disruption cases, including feasible and infeasible instances, and 200 feasible simultaneous compound-disruption cases. We compare LLM-Z3 full replanning, IPyHOPPER hierarchical repair, and an iTIMO local-revision adapter across effectiveness, plan stability, and computational cost. LLM-Z3 with Gemini achieved the highest observed compound-disruption success. IPyHOPPER nearly matched that configuration's single-disruption overall success, while preserving substantially more of the accepted itinerary on successful repairs. Successful hierarchical and local repairs made fewer edits and retained more accepted commitments than full replanning. Computational profiles differed: IPyHOPPER used no LLM inference, the evaluated LLM-Z3 adapter used compact one-call inference, and the iTIMO adapter consumed substantially more tokens. The study provides practical guidelines for balancing feasibility recovery, commitment preservation, and computational cost within evaluated settings.

Every Monday
Get next week’s papers.
Subscribe on Substack