Research note

SuperGPQA: Graduate-Level Knowledge at Frontier Scale

2026-06-28 · SuperGPQA · 1 min read

26,529
questions
285
graduate subfields
9.67
average answer options
61.82
top sample accuracy

How far does frontier-model knowledge really extend beyond the mainstream disciplines? We examine graduate-level question answering across hundreds of specialized fields and find that long-tail expertise remains the clearest gap between models.

The mainstream illusion

Most knowledge benchmarks sample from the same comfortable center: mathematics, physics, computer science, a slice of law and medicine. Models ace them, leaderboards saturate, and it becomes tempting to conclude that knowledge is solved. SuperGPQA is built on the opposite suspicion — that the center was never the hard part.

The benchmark spans 285 graduate-level disciplines, reaching past the familiar STEM core into fields most evaluation suites have never touched. Questions are filtered through a human–LLM collaborative pipeline that removes trivial or ambiguous items, with large-scale expert annotation keeping the bar at genuine graduate difficulty.

What the leaderboard says

Scores are sample-level overall accuracy on the main leaderboard:

#ModelOverall (sample)
1DeepSeek-R161.82
2o1-2024-12-1760.24
2DeepSeek-R1-Zero60.24

Two things stand out. First, even the strongest reasoning models sit near 60 — on a 100-point scale, with every discipline in play, nobody is close to done. Second, the ordering compresses: models that look far apart on mainstream suites land within a few points of each other here, because no amount of instruction tuning substitutes for coverage of the long tail.

Why the long tail matters

A model that answers convincingly in twenty popular fields and quietly fabricates in the other two hundred is not a knowledge engine; it is a demo. The practical value of an assistant in materials science, veterinary pathology, or maritime law lives precisely where benchmark coverage has historically been thinnest. SuperGPQA makes that thin region measurable — and once measurable, improvable.

Beyond the scores, the paper's discipline-level methodology doubles as a blueprint for anyone who needs to build wide-coverage evaluations without letting quality collapse at scale.