MM-BrowseComp: Multimodal Browsing Beyond Text Trails
Web research is not text-only. MM-BrowseComp asks search agents to follow evidence through images, charts, and pages that no text-only pipeline can shortcut, revealing a wide gap behind text-centric browsing scores.
The web keeps its best secrets in pixels. Product details live in screenshots, scientific claims in figures, historical facts in scanned documents. MM-BrowseComp builds browsing tasks whose evidence chain runs through images, charts, and video — so an agent that only reads will follow the trail to a dead end.
The gap nobody's text benchmark shows
The Overall column is not a ranking so much as a cliff:
| # | Model | Overall Accuracy (Pass@1) |
|---|---|---|
| 1 | o3 (official tools) | 29.02 |
| 2 | Gemini 2.5 Pro Preview 05-06 (official tools) | 7.14 |
| 3 | Gemini 2.5 Flash Preview 05-20 (official tools) | 3.12 |
One system reaches 29; the next lands below 8. A twenty-one point drop between first and second place is the kind of result that usually signals a capability boundary rather than incremental engineering — most of today's agents browse with their eyes closed, and this benchmark is the first mirror that shows it.
Why the tasks resist shortcuts
- Questions require joint reasoning over text, images, and video encountered mid-trajectory — not just at the final page.
- Key hops are deliberately un-searchable as plain text, so query rewriting alone cannot rescue a blind agent.
- Verification steps punish confident fabrication: the right answer needs the right visual evidence, found live.
We expect this table to look very different within a year. Multimodal tool use is improving quickly, and MM-BrowseComp is positioned to be the benchmark where that progress first becomes visible — or embarrassingly, where it fails to.
Task construction details and the full agent table are on the MM-BrowseComp benchmark page.
MM-BrowseComp — Multimodal Browsing · Agent Benchmark