CodeEditorBench
Execution-checked edits across debugging, translation, optimization, and changing requirements.
Overview
CodeEditorBench starts from an observation that code generation benchmarks miss much of day-to-day maintenance. A model may need to repair a faulty program, translate working logic into another language, make an implementation faster without changing its behavior, or adapt existing code after the requirement changes. The benchmark turns those four situations into 7,961 execution-checked tasks in C++, Java, and Python.
The tasks are sourced from LeetCode, CodeContests, CodeXGLUE, CodeNet, and TACO. They carry an average of 44 tests, with no task having fewer than eight and some having as many as 446. Missing and boundary inputs are supplemented with GPT-4, GLM-4, and Qwen-72B-Chat, but expected outputs come from executing the reference program. Generated edits are cleaned, placed into language-specific templates, compiled, and run in a Dockerized online judge.
The paper distinguishes the original Primary split from CodeEditorBench_Plus, which applies timestamp filtering to reduce contamination. Table 1 reports both values in each cell: Plus outside parentheses and Primary inside. TokenWave uses only the Plus, zero-shot, greedy-decoding Win Rate so every displayed model shares the same split, prompting regime, decoding strategy, and rank aggregation.
Why it matters
Editing exposes constraints that a fresh-generation prompt can avoid. The answer must preserve behavior that is still correct, satisfy a target language or changed specification, fit an existing I/O wrapper, and survive a larger test suite. For polishing, functional correctness is only the entry condition; the edit must also reduce measured runtime or memory. These requirements make a high score evidence of controlled modification rather than code-shaped text generation.
The Plus split matters because the paper observes differences above 0.2 between Primary and Plus pass@1 for some model-category pairs. A model can appear stronger when older source problems are familiar. Keeping the timestamp-filtered values separate makes the headline comparison less vulnerable to that effect.
Contributions
- Defines four execution-checked code-editing scenarios—debugging, translation, polishing, and requirement switching—rather than treating editing as ordinary code completion.
- Releases 7,961 tasks sourced from LeetCode, CodeContests, CodeXGLUE, CodeNet, and TACO, with an average of 44 tests and explicit difficulty, language, error-count, and relation-strength annotations.
- Introduces CodeEditorBench_Plus, a timestamp-filtered split intended to reduce contamination visible in the original Primary split.
- Provides a Dockerized online-judge workflow, task-specific pass criteria, zero-shot and three-shot prompts, and a rank-normalized Win Rate for comparing performance across heterogeneous editing tasks.
Method & evaluation
- Source programs longer than 800 lines or roughly 1,000 tokens are filtered out. The retained tasks cover C++, Java, and Python and are stratified by language, easy/medium/hard difficulty, injected error count, translation direction, and strong or weak requirement relation.
- Debug tasks inject one to four error types. Translation and polishing examples are sampled by code complexity in a 3:4:1 ratio. Requirement-switch tasks pair closely related LeetCode problems or weakly related problems found by shared tags and BERT similarity of at least 0.92.
- GPT-4, GLM-4, and Qwen-72B-Chat supplement missing and boundary test inputs; expected outputs are produced by running the reference solution in the online judge. Code templates add the headers, I/O handling, and wrappers required for local compilation.
- The paper evaluates 19 models with greedy decoding in zero-shot and three-shot settings. Maximum new tokens are 1,024 for API models and 2,048 for open models; generated text is cleaned, parsed, integrated into a language template, and compiled before judging.
- For Debug, Translate, and Requirement Switch, a task passes only when all tests succeed under time and memory limits. For Polish, the original runtime and memory are averaged over 20 executions; a correct generated edit is measured twice and must reduce runtime or memory to receive a non-zero optimization score.
- The displayed leaderboard takes the Plus value outside parentheses from the zero-shot block of Table 1. Per-category ranks are converted to 1 − (rank − 1) / 19 and averaged to obtain Win Rate.
Evaluation metrics
Plus zero-shot Win Rate — Higher is better. Mean rank-normalized score across Debug, Translate, Switch, and Polish on CodeEditorBench_Plus; each category contributes 1 − (rank − 1) / 19. Score range: [0, 1].
pass@1 — Higher is better. Fraction of greedy model outputs that pass every online-judge test for Debug, Translate, or Requirement Switch; translation is executed in the target-language environment. Score range: [0, 1].
Mean OptScore — Higher is better. Mean non-negative relative reduction in runtime and memory for Polish outputs that remain functionally correct; incorrect programs receive zero. Score range: [0, 1].
Leaderboard
Top six rows from Table 1 under the same setting: CodeEditorBench_Plus values (outside parentheses), zero-shot prompting, and greedy decoding. Win Rate is an average of relative category ranks, not a raw task pass percentage.
| # | Model | Plus zero-shot Win Rate |
|---|---|---|
| 1 | GPT-4 | 0.855 |
| 2 | OpenCI-DS-33B | 0.776 |
| 3 | Gemini-Ultra | 0.75 |
| 4 | DS-33B-INST | 0.737 |
| 5 | Gemini-Pro | 0.737 |
| 6 | GPT-3.5-Turbo | 0.724 |
Reading the results
GPT-4 leads the Plus zero-shot table at 0.855. OpenCI-DS-33B is second at 0.776, a gap of 0.079, and is the strongest open model in this setting. Gemini-Ultra follows at 0.750. DS-33B-INST and Gemini-Pro tie at 0.737, while GPT-3.5-Turbo closes the displayed top six at 0.724.
Win Rate is not the fraction of all 7,961 tasks solved. Each model is ranked separately on Debug, Translate, Switch, and Polish; a rank is converted to 1 − (rank − 1) / 19, then the four values are averaged. GPT-4's 0.855 therefore means consistently strong relative placement across task types, while a model with one exceptional specialty and several weak categories can rank lower despite a comparable raw pass rate in that specialty.
The category distribution explains why one aggregate is not enough. Across model outputs on Plus, requirement switching has an 11.18% pass rate, debugging is about 20%, and translation about 30%. For polishing, 37.47% of outputs remain functionally correct, but only 19.23% both pass the tests and improve runtime or memory. The benchmark's hardest failures are thus not confined to syntax: preserving intent under a changed requirement and producing a real optimization remain uncommon.
Figures & tables