SuperGPQA: Graduate-Level Knowledge at Frontier Scale
How far does frontier-model knowledge really extend beyond the mainstream disciplines? We examine graduate-level question answering across hundreds of specialized fields and find that long-tail expertise remains the clearest gap between models.
The mainstream illusion
Most knowledge benchmarks sample from the same comfortable center: mathematics, physics, computer science, a slice of law and medicine. Models ace them, leaderboards saturate, and it becomes tempting to conclude that knowledge is solved. SuperGPQA is built on the opposite suspicion — that the center was never the hard part.
The benchmark spans 285 graduate-level disciplines, reaching past the familiar STEM core into fields most evaluation suites have never touched. Questions are filtered through a human–LLM collaborative pipeline that removes trivial or ambiguous items, with large-scale expert annotation keeping the bar at genuine graduate difficulty.
What the leaderboard says
Scores are sample-level overall accuracy on the main leaderboard:
| # | Model | Overall (sample) |
|---|---|---|
| 1 | DeepSeek-R1 | 61.82 |
| 2 | o1-2024-12-17 | 60.24 |
| 2 | DeepSeek-R1-Zero | 60.24 |
Two things stand out. First, even the strongest reasoning models sit near 60 — on a 100-point scale, with every discipline in play, nobody is close to done. Second, the ordering compresses: models that look far apart on mainstream suites land within a few points of each other here, because no amount of instruction tuning substitutes for coverage of the long tail.
Why the long tail matters
A model that answers convincingly in twenty popular fields and quietly fabricates in the other two hundred is not a knowledge engine; it is a demo. The practical value of an assistant in materials science, veterinary pathology, or maritime law lives precisely where benchmark coverage has historically been thinnest. SuperGPQA makes that thin region measurable — and once measurable, improvable.
Beyond the scores, the paper's discipline-level methodology doubles as a blueprint for anyone who needs to build wide-coverage evaluations without letting quality collapse at scale.
Explore the full leaderboard, metric definitions, and evaluation setup on the SuperGPQA benchmark page.
SuperGPQA — Graduate Knowledge · LLM Benchmark