Research · Graduate Knowledge

SuperGPQA

Graduate multiple-choice evaluation across 285 subfields, with sample-weighted and taxonomy-balanced views.

Benchmark paper · Leader — DeepSeek-R1 · 61.82 · Graduate Knowledge · updated 2026-08-04

26,529
questions
285
graduate subfields
9.67
average answer options
61.82
top sample accuracy

Overview

SuperGPQA is a 26,529-question multiple-choice benchmark organized into 13 disciplines, 72 fields, and 285 graduate subfields. Every subfield contributes at least 50 questions, and the average item has 9.67 options rather than the conventional four. The benchmark includes long-tail areas such as agriculture, military science, light-industry engineering, and service-oriented disciplines alongside science, engineering, and medicine; 42.33% of all questions require calculation or formal reasoning.

Construction is a staged human–LLM process. Expert annotators first collect questions and solutions from credible sources. Crowd annotators translate, convert non-choice problems, standardize complex statement-selection items, and generate distractors. Rule checks validate the schema, multiple contemporary models flag suspicious or low-discrimination questions, and experts then solve and revise flagged items with unrestricted web access. The paper reports more than 80 expert annotators and more than 30 minutes of review per suspicious candidate in post-annotation interviews.

Question counts are intentionally not uniform: Science, Engineering, and Medicine make up 77.2% of samples because more items in those areas survived the difficulty and quality filters. SuperGPQA therefore reports four overall views. Overall (sample) treats every question equally, while Overall (subfield), Overall (field), and Overall (discipline) average equally within the corresponding taxonomy level. This page ranks the zero-shot reasoning and chat models by the first metric and keeps the macro scores separate.

Why it matters

Broad subject labels can hide narrow coverage. A model may look strong on 'science' while failing specialized optics, metallurgy, veterinary medicine, or transportation engineering. The 285-subfield taxonomy allows failures to be localized at the level where professional use actually occurs, while the minimum question count preserves enough samples for each subfield to be measured.

The multiple aggregation levels prevent one leaderboard number from silently inheriting the dataset's STEM concentration. Sample accuracy is useful for the full released question pool; discipline-macro accuracy asks whether performance is balanced across thirteen top-level areas. Comparing the two reveals whether a system's gains come from the largest categories or transfer to the long tail.

Contributions

  • Builds a 26,529-question graduate benchmark spanning 13 disciplines, 72 fields, and 285 subfields, with at least 50 questions per subfield and an average of 9.67 answer options.
  • Develops a large-scale human–LLM construction workflow covering expert source screening, translation and multiple-choice conversion, distractor generation, plagiarism checks, and rule/model/expert quality inspection.
  • Reports sample-weighted and macro-averaged accuracy at three taxonomy levels, plus Easy, Middle, and Hard splits, so performance is not represented only by the STEM-heavy sample distribution.
  • Evaluates reasoning, chat, instruct, and base models under disclosed zero-shot or five-shot protocols and analyzes prompt robustness, subfield cues, scaling, difficulty profiles, and disciplinary discrimination.

Method & evaluation

  • Expert annotators collect questions and solutions from credible sources; crowd annotators translate non-English material, convert open questions to multiple choice, standardize statement-combination items, and add plausible distractors.
  • Rule checks validate formatting, options, answers, difficulty, and taxonomy fields. Seven contemporary LLMs flag invalid, ambiguous, trivial, multimodal, incomplete, or low-discrimination items for unrestricted expert review, which averages more than 30 minutes per suspicious question.
  • The final 26,529 items cover 285 subfields with at least 50 each. Science, Engineering, and Medicine contribute 77.2% of samples, so the paper reports both sample-weighted accuracy and unweighted averages across subfields, fields, and disciplines.
  • Reasoning and chat models are evaluated zero-shot; base models use a five-shot protocol following MMLU-Pro. Temperature is 0, maximum generation is 32k tokens for reasoning models and 4k for all other models.
  • Exact-match multiple-choice accuracy is computed overall and by Easy, Middle, and Hard labels. Additional experiments vary the presence of subfield names and evaluate Qwen2.5 models under 24 semantically equivalent prompt formats.

Evaluation metrics

Overall (sample) — Higher is better. Micro-averaged multiple-choice accuracy across all 26,529 questions; this is the metric used by the leaderboard. Score range: [0, 100].

Overall (subfield) — Higher is better. Unweighted mean of the 285 subfield accuracies, reducing the influence of subfields with more questions. Score range: [0, 100].

Overall (field) — Higher is better. Unweighted mean accuracy across 72 fields. Score range: [0, 100].

Overall (discipline) — Higher is better. Unweighted mean accuracy across 13 top-level disciplines, the broadest macro-balanced view. Score range: [0, 100].

Leaderboard

The ranking uses sample-weighted accuracy from the paper's zero-shot reasoning and chat-model setting. Five-shot base-model rows are not mixed into this table; macro subfield, field, and discipline scores remain available as separate metrics.

#ModelOverall (sample)
1DeepSeek-R161.82
2o1-2024-12-1760.24
2DeepSeek-R1-Zero60.24
4o3-mini-2025-01-31-high55.22
5Doubao-1.5-Pro-32k-25011555.09
6o3-mini-2025-01-31-medium52.69
7Doubao-1.5-Pro-32k-24122550.93
8Qwen-Max-2025-01-2550.08
9Claude 3.5 Sonnet-2024102248.16
10o3-mini-2025-01-31-low48.03

Reading the results

DeepSeek-R1 leads zero-shot sample accuracy at 61.82%. o1-2024-12-17 and DeepSeek-R1-Zero tie at 60.24%, a detail missing from the previous page. A second cluster begins with o3-mini-high at 55.22 and Doubao-1.5-Pro-250115 at 55.09, followed by o3-mini-medium at 52.69. The top ten shown here all share the zero-shot reasoning/chat protocol; the paper's five-shot base-model results are not inserted into the same ranking.

Difficulty splits show why the overall order alone is incomplete. DeepSeek-R1 scores 63.59 on Easy, 63.63 on Middle, and 56.87 on Hard. o3-mini-high moves in the opposite direction—53.05, 56.09, and 56.16—supporting the paper's distinction between broad professional recall and reasoning-focused behavior. Macro scores can also change the leader: DeepSeek-R1-Zero posts 60.99 at the discipline level, above DeepSeek-R1's 59.95, despite trailing it on sample-weighted accuracy.

Figures & tables

Figure 2 (paper, PDF p. 6): SuperGPQA's source-screening, transcription, and quality-inspection pipeline, including expert sourcing, distractor generation, and rule-, model-, and human-based checks. Source: paper.
Figure 2 (paper, PDF p. 6): SuperGPQA's source-screening, transcription, and quality-inspection pipeline, including expert sourcing, distractor generation, and rule-, model-, and human-based checks. Source: paper.
Figure 7 (paper, PDF p. 20): Qwen2.5 accuracy from 0.5B to 72B with and without subfield cues, and robustness across 24 semantically equivalent prompts; larger models gain more from cues and show lower prompt variance. Source: paper.
Figure 7 (paper, PDF p. 20): Qwen2.5 accuracy from 0.5B to 72B with and without subfield cues, and robustness across 24 semantically equivalent prompts; larger models gain more from cues and show lower prompt variance. Source: paper.