CodeSimpleQA
Bilingual factual programming questions, scored separately from executable code generation.
Overview
CodeSimpleQA tests a layer of coding competence that execution benchmarks do not isolate: whether a model can state a programming fact correctly. Its 1,498 questions comprise 782 English and 716 Chinese items, span more than 15 programming languages and 21 computer-science domains, and use concise reference answers grounded in code-related documents. The distribution is broad rather than uniform: general-language questions dominate, while individual language subsets include Java, Python, C#, C/C++, SQL, PHP, JavaScript, Kotlin, Ruby, Swift, Go, Rust, and others.
The evaluation follows the SimpleQA distinction among correct, incorrect, and not-attempted responses. Correct Given Attempted measures how often a model is right after it chooses to answer; F-score is the harmonic mean of that value and the overall Correct rate. This matters because a system that guesses constantly and one that refuses difficult questions can otherwise look deceptively similar under a single accuracy number.
The paper also treats the benchmark as a target for model improvement. CodeSimpleQA-Instruct contains 53,571,094 English and 13,359,625 Chinese samples. Its pipeline retrieves technical documents, filters source and content quality, clusters code-text embeddings, generates objective QA pairs, and applies automated plus human verification before SFT and GRPO experiments. The benchmark and training corpus are related, but the leaderboard here uses only held-out CodeSimpleQA evaluation results.
Why it matters
Executable code can still be built on a false premise: a hallucinated API contract, a mistaken language rule, or an incorrect systems claim. Those errors surface in code review, debugging, migration, and technical support, where an authoritative-sounding answer may redirect the implementation before any test is run. CodeSimpleQA makes that factual layer visible instead of allowing execution success to stand in for knowledge.
The bilingual split also prevents an English result from hiding uneven transfer. TokenWave therefore averages the paper's Chinese and English overall F-scores for models present in both tables. The combined number is deliberately labeled as a site-derived mean; readers who need language-specific or domain-specific diagnosis should use the two source tables and the paper's 21-domain breakdown.
Contributions
- Introduces a bilingual factual-code benchmark with 1,498 items: 782 English questions and 716 Chinese questions, all paired with concise reference answers.
- Covers more than 15 programming languages and 21 computer-science domains, with answers grounded in technical documents rather than inferred from whether generated code executes.
- Uses the SimpleQA-style correct, incorrect, and not-attempted outcome taxonomy and reports Correct Given Attempted and F-score so abstention is not treated the same as a factual error.
- Releases CodeSimpleQA-Instruct with 53,571,094 English and 13,359,625 Chinese samples and evaluates factuality-oriented SFT and GRPO post-training.
Method & evaluation
- The evaluation set contains 782 English and 716 Chinese QA pairs distributed across 21 domains; software engineering, web technologies, programming languages, and operating systems are among the largest categories.
- Questions are drawn from Code WebQA, self-constructed material, and human-written content grounded in sources such as official documentation, GitHub, and Stack Overflow. Eight annotators and three engineers curate the data; a manually constructed subset retains 312 of roughly 1,500 candidates after review by at least three people.
- Models receive a short factual question and are instructed to answer in no more than 64 words. A reference-aware judge assigns CORRECT, INCORRECT, or NOT_ATTEMPTED according to whether the response contains the reference answer and whether it introduces a contradiction.
- The paper reports domain-level and overall F-scores separately for Chinese and English. The TokenWave table uses only models appearing in both tables and takes the arithmetic mean of the two reported overall values.
- The companion training pipeline recalls code-related Common Crawl documents, filters for source and content quality, clusters code-text embeddings with DBSCAN, generates QA pairs at low temperature, and retains only pairs that pass automated and human quality checks.
Evaluation metrics
Bilingual mean F1 — Higher is better. TokenWave-derived arithmetic mean of the Chinese Avg. and English Avg. F-scores reported by the paper; used only for this site's bilingual leaderboard. Score range: [0, 100].
F-score — Higher is better. Harmonic mean of the Correct rate and Correct Given Attempted, reported by domain and as an overall average for each language. Score range: [0, 100].
Correct — Higher is better. Percentage of responses that fully contain the reference answer without contradictory content. Score range: [0, 100].
Correct Given Attempted — Higher is better. Percentage of correct responses among questions the model attempted; separates answer precision from willingness to answer. Score range: [0, 100].
Leaderboard
Each score is the arithmetic mean of the paper's Chinese Avg. and English Avg. F-scores. This site aggregation makes bilingual performance readable in one table; it is not a third metric introduced by the authors.
| # | Model | Bilingual mean F1 |
|---|---|---|
| 1 | GPT-5 | 65.05 |
| 2 | o3 | 61.65 |
| 3 | GPT-5 Mini | 55.65 |
| 4 | o4-mini | 53.3 |
| 5 | Claude Sonnet 4 Thinking (2025-05-14) | 52.5 |
| 6 | Claude Sonnet 4 (2025-05-14) | 51.9 |
| 7 | ChatGPT-4o Latest | 50.95 |
| 8 | o3-mini-2025-01-31 | 50.45 |
| 9 | GPT-OSS-120B | 50.35 |
| 10 | GPT-4o-2024-11-20 | 48.7 |
Reading the results
GPT-5 leads the bilingual aggregation at 65.05, formed from 67.2 Chinese F-score and 62.9 English F-score. o3 follows at 61.65, a 3.40-point gap, while GPT-5 Mini is another 6.00 points back at 55.65. The first separation is therefore between the two strongest reasoning-oriented systems and the rest of the field, not merely between vendors.
The middle of the table is compact: o4-mini scores 53.30, Claude Sonnet 4 Thinking 52.50, and Claude Sonnet 4 without thinking 51.90. ChatGPT-4o Latest, o3-mini, and GPT-OSS-120B all fall between 50.35 and 50.95. Small changes in answer rate or calibration can reorder this band, so the per-language Correct and Correct Given Attempted metrics are useful when comparing nearby models.
GPT-4o-2024-11-20 closes the displayed top ten at 48.70, leaving a 16.35-point spread from first to tenth. More importantly, the leader remains 34.95 points below the top of the scale. Even on short, reference-grounded programming questions, the paper's results do not support treating factual code knowledge as solved.
Figures & tables