Research note

MM-BrowseComp: Multimodal Browsing Beyond Text Trails

2026-03-30 · MM-BrowseComp · 1 min read

224
verified questions
22
multimodal subtasks
57%
prompts with images
29.02
top tool-augmented OA

Web research is not text-only. MM-BrowseComp asks search agents to follow evidence through images, charts, and pages that no text-only pipeline can shortcut, revealing a wide gap behind text-centric browsing scores.

The web keeps its best secrets in pixels. Product details live in screenshots, scientific claims in figures, historical facts in scanned documents. MM-BrowseComp builds browsing tasks whose evidence chain runs through images, charts, and video — so an agent that only reads will follow the trail to a dead end.

The gap nobody's text benchmark shows

The Overall column is not a ranking so much as a cliff:

#ModelOverall Accuracy (Pass@1)
1o3 (official tools)29.02
2Gemini 2.5 Pro Preview 05-06 (official tools)7.14
3Gemini 2.5 Flash Preview 05-20 (official tools)3.12

One system reaches 29; the next lands below 8. A twenty-one point drop between first and second place is the kind of result that usually signals a capability boundary rather than incremental engineering — most of today's agents browse with their eyes closed, and this benchmark is the first mirror that shows it.

Why the tasks resist shortcuts

  • Questions require joint reasoning over text, images, and video encountered mid-trajectory — not just at the final page.
  • Key hops are deliberately un-searchable as plain text, so query rewriting alone cannot rescue a blind agent.
  • Verification steps punish confident fabrication: the right answer needs the right visual evidence, found live.

We expect this table to look very different within a year. Multimodal tool use is improving quickly, and MM-BrowseComp is positioned to be the benchmark where that progress first becomes visible — or embarrassingly, where it fails to.