MM-BrowseComp
Open-web questions whose decisive evidence lives in images and video, with the reasoning path checked step by step.
Overview
MM-BrowseComp targets questions that cannot be solved by reading only webpage text. Its 224 items span 22 subtasks in media, technology, society, geography, and academics. Some prompts begin with an image and others begin with text, but all require retrieving and reasoning over image or video evidence somewhere in the browsing process. The final answers remain short phrases—names, numbers, or colors—so answer checking is tractable even though the search path is open-ended.
Every question also carries an irreducible checklist: an ordered set of evidence and reasoning steps that should be necessary to reach the answer. This supports three levels of scoring. Overall Accuracy checks the answer alone; Strict Accuracy requires both the correct answer and completion of every checklist item; Average Checklist Score gives partial credit for completing portions of the path. A correct guess can therefore raise OA without being counted as a strict solution.
The construction process started with 300 candidates authored by more than twenty master's- and PhD-level researchers. Questions had to resist a single web-enabled attempt by Gemini 2.5 Pro and GPT-4o, remain unsolved by another annotator after five minutes of active search, avoid text-only shortcuts, and preserve a unique, temporally stable answer. After pilot calibration, secondary review, tool-dependency checks, and factual verification, 161 items were accepted directly, 63 revised, and 76 discarded.
Why it matters
A browsing agent can fail before reasoning begins: it may never open the relevant image, lack a video-processing tool, or discard visual details when handing work to a text-only subagent. Final-answer scoring collapses all of those failures into one zero. MM-BrowseComp's checklists expose where the pipeline breaks and whether an apparently correct answer is actually supported by the retrieved evidence.
The benchmark also distinguishes native multimodal tool use from strong language reasoning wrapped around lossy visual tools. That distinction is operationally important for research systems: adding more independent attempts can improve the chance of landing on the right answer, but the paper's scaling experiment shows little corresponding gain in Strict Accuracy, indicating that repeated sampling does not repair the underlying reasoning path.
Contributions
- Introduces 224 difficult multimodal web-search questions across 22 subtasks and five categories; 57% have image-bearing prompts and all require image or video evidence during retrieval.
- Associates every question with a concise, temporally stable answer and an irreducible checklist that specifies the minimum evidence and reasoning path needed to derive it.
- Uses a three-phase construction and validation process that filters questions solvable by Gemini 2.5 Pro or GPT-4o without tools, by a second annotator within five minutes, or by simple textual shortcuts.
- Separates final-answer accuracy from reasoning-path validity through Overall Accuracy, Strict Accuracy, and Average Checklist Score, including modality-specific checklist analysis and test-time scaling experiments.
Method & evaluation
- More than twenty master's- and PhD-level AI researchers authored questions across 22 assigned subtasks; the five top-level categories account for 29% Media, 26% Technology, 18% Society, 13% Geography, and 14% Academics.
- Annotators begin from a known fact and reverse-engineer a multi-hop question whose answer is short, unique, and stable. Essential evidence must live in images or video rather than be recoverable from a text-only shortcut.
- Each item includes an ordered, irreducible checklist. The final pool was reduced from 300 candidates to 224: 161 accepted directly, 63 revised, and 76 discarded after calibration, secondary review, tool-dependency screening, and factual verification.
- The main table separates tool-free VLMs, official tool-augmented VLMs, and open-source agent frameworks. Tool-free and official tool-augmented systems use all 224 questions; compute-intensive open-source agents use a stratified 54-question subset and are not mixed into this page's leaderboard.
- All results are Pass@1. GPT-4o-2024-11-20 judges final-answer correctness and checklist completion from a required answer plus problem-solving roadmap; Strict Accuracy passes only when both the answer and every checklist item pass.
Evaluation metrics
Overall Accuracy (OA) — Higher is better. Pass@1 percentage of questions with a correct final answer, independent of whether the full checklist is demonstrated. Score range: [0, 100].
Strict Accuracy (SA) — Higher is better. Pass@1 percentage for which the final answer is correct and every item in the associated irreducible reasoning checklist is completed. Score range: [0, 100].
Average Checklist Score (AVG CS) — Higher is better. Mean fraction of checklist items completed across questions, providing partial credit for valid intermediate retrieval and reasoning steps. Score range: [0, 100].
Leaderboard
This table keeps only the paper's full 224-question Tool-Augmented VLM setting. Tool-free VLMs use a different tool condition, and open-source agents were evaluated on a 54-question subset, so neither group is mixed into the ranking.
| # | Model | Overall Accuracy (Pass@1) |
|---|---|---|
| 1 | o3 (official tools) | 29.02 |
| 2 | Gemini 2.5 Pro Preview 05-06 (official tools) | 7.14 |
| 3 | Gemini 2.5 Flash Preview 05-20 (official tools) | 3.12 |
Reading the results
Within the comparable full-dataset tool-augmented group, o3 leads at 29.02% Overall Accuracy. Gemini 2.5 Pro Preview 05-06 reaches 7.14% and Gemini 2.5 Flash Preview 05-20 reaches 3.12%, leaving a 21.88-point gap between first and second. The tool-free and open-source-agent rows in the paper use different tool or dataset conditions and are deliberately excluded from this ranking.
The diagnostic metrics are less forgiving than final answers. o3 records 19.64% Strict Accuracy and 36.49% Average Checklist Score, compared with 29.02% OA. Its modality-specific checklist completion is 62.13% on textual items and 52.72% on image/video items. For Agent-R1, increasing independent runs from one to sixteen lifts OA under aggregation, while SA remains around the low single digits; the extra compute increases successful answer sampling more than it improves complete multimodal reasoning.
Figures & tables
Read the research note
MM-BrowseComp: Multimodal Browsing Beyond Text Trails