Research · Multimodal Planning

WorldTravel

Itineraries checked for executable temporal feasibility after agents recover constraints from rendered travel pages.

Benchmark paper · Leader — GPT-5.2 · 19.33 · Multimodal Planning · updated 2026-08-04

150
travel scenarios
2,003
rendered webpages
15+
constraints per task on average
19.33
top multimodal feasibility

Overview

WorldTravel treats itinerary generation as a constrained scheduling problem rather than a prose task. The benchmark contains 150 scenarios, with 30 each in Berlin, Vienna, Rome, Barcelona, and Florence. Each task requires coordinating between 10 and 23 hard and soft constraints, more than 15 on average, and up to five temporal anchors such as timed museum entries, performances, or restaurant reservations. A single invalid slot, closure conflict, short dwell time, or missed travel buffer makes the full itinerary infeasible.

WorldTravel-Webscape supplies the information through 2,003 rendered webpages covering 36 attractions, 25 restaurants, 26 hotels, booking systems, menus, guides, and transportation matrices. Eight environment APIs return screenshots rather than JSON records. Relevant facts are distributed across page types and visual states—such as a grayed-out sold-out slot—so the agent must retrieve, perceive, and reconcile parameters before planning.

Evaluation is executable and decomposed. Feasibility Rate requires every hard temporal check to pass. Constraint Violation measures the fraction of hard constraints missed even when the complete plan fails. Optimality given Feasible evaluates cost calculations and preference choices only among feasible itineraries. The main table includes both text and multimodal conditions; this page keeps the seven multimodal rows together because they share the screenshot-returning environment.

Why it matters

Travel planning exposes two independent failure surfaces. An agent can misread the environment—confusing availability, hours, prices, or dwell-time guidance—or it can extract every parameter and still fail to coordinate them globally. The paired text and screenshot settings estimate the first gap, while the Gold-parameter ablation shows that perfect access does not remove the second.

The all-hard-constraints definition avoids the common problem of fluent but unusable plans receiving high quality scores. It also generalizes beyond travel: production scheduling, logistics, appointment coordination, and resource allocation all share the property that local choices consume options downstream and one violated temporal boundary can invalidate the whole output.

Contributions

  • Introduces 150 automatically verifiable travel-planning scenarios across five European cities, each combining 10–23 temporal, cost, and preference constraints with more than 15 constraints on average.
  • Builds WorldTravel-Webscape with 2,003 rendered webpages covering 36 attractions, 25 restaurants, 26 hotels, booking interfaces, guides, menus, and transportation matrices.
  • Formalizes four hard temporal constraint classes and two soft decision classes, with task-specific boolean verification functions for feasibility, violation rate, and conditional optimality.
  • Evaluates ten text models and seven multimodal models, isolating a visual Perception–Action Gap and a planning-capacity threshold near ten hard constraints through text, screenshot, and Gold-parameter ablations.

Method & evaluation

  • The benchmark contains 30 tasks for each of Berlin, Vienna, Rome, Barcelona, and Florence. Experts combine official operating and pricing data with practical dwell-time, queue, and travel information, then adjust availability to create tightly coupled but solvable itineraries.
  • WorldTravel-Webscape distributes facts across 2,003 static pages with varied layouts. Its eight attraction, restaurant, hotel, and routing APIs return rendered screenshots, forcing OCR, UI-state interpretation, and cross-page evidence integration.
  • Every itinerary is represented as ordered activities with start times, durations, and discrete choices. Hard checks cover timed-entry slots, operating windows, minimum dwell time, and travel plus arrival buffers; soft checks cover exact cost combinations and hotel selection.
  • Each task passes author review, independent cross-solving, and external validation, and includes a reference itinerary, explicit constraint annotations, and executable verification functions. Pilot runs filter ambiguous or trivial cases.
  • The main experiment evaluates 150 tasks in a text setting with parameters extracted into text and a multimodal setting with screenshot-returning tools. The leaderboard uses only the latter; a 30-task Gold-parameter ablation separately tests planning under perfect hard-constraint access.

Evaluation metrics

Feasibility Rate — Higher is better. Percentage of tasks for which every hard temporal constraint is satisfied; one hard violation makes the complete itinerary infeasible. Score range: [0, 100].

Constraint Violation — Lower is better. Average percentage of hard constraints violated across all tasks, retaining partial diagnostic signal when the itinerary is not fully feasible. Score range: [0, 100].

Optimality | Feasible — Higher is better. Average soft-constraint satisfaction for cost computation and preference choices, calculated only over temporally feasible itineraries. Score range: [0, 100].

Leaderboard

All seven rows use the 150-task multimodal setting in which the environment APIs return rendered webpage screenshots. Text-setting scores and the 30-task Gold-parameter ablation are separate conditions and are not mixed into this ranking.

#ModelFeasibility Rate (%)
1GPT-5.219.33
2Claude Opus 4.514
3Gemini 3 Pro8
4Doubao-1.8-Pro4
5GPT-5.12.67
5Claude Sonnet 4.52.67
7Gemini 2.5 Pro1.33

Reading the results

GPT-5.2 leads multimodal Feasibility Rate at 19.33%, followed by Claude Opus 4.5 at 14.00% and Gemini 3 Pro at 8.00%. Doubao-1.8-Pro reaches 4.00%; GPT-5.1 and Claude Sonnet 4.5 tie at 2.67%; Gemini 2.5 Pro records 1.33%. These are full-task pass rates, so a system can satisfy many individual constraints and still receive zero feasibility on a scenario.

The paired settings quantify the perception cost. GPT-5.2 falls from 32.67% with text-extracted parameters to 19.33% from screenshots, while Claude Opus 4.5 falls from 21.33% to 14.00% and Gemini 3 Pro from 14.67% to 8.00%. Planning remains a separate bottleneck: across text models, average feasibility drops from 36.3% at 6–7 hard constraints to 18.3% at 8–9, 3.4% at 10–11, and 3.1% at 12 or more. Timed-entry and operating-window constraints average only 27% satisfaction in text and 16% in vision.

Figures & tables

Figure 2 (paper, PDF p. 3): Data collection, city selection, webpage synthesis, expert task design, and manual plus multi-model quality control that produces 150 verified tasks. Source: paper.
Figure 2 (paper, PDF p. 3): Data collection, city selection, webpage synthesis, expert task design, and manual plus multi-model quality control that produces 150 verified tasks. Source: paper.
Figure 5 (paper, PDF p. 10): Text-setting feasibility as hard constraints increase and average satisfaction by constraint type in text versus vision; feasibility drops sharply at 10 constraints and timed-entry slots are hardest. Source: paper.
Figure 5 (paper, PDF p. 10): Text-setting feasibility as hard constraints increase and average satisfaction by constraint type in text versus vision; feasibility drops sharply at 10 constraints and timed-entry slots are hardest. Source: paper.