Research · Repository Generation

NL2Repo-Bench

Complete Python libraries generated from one requirements document, with correctness decided by the original hidden test suite.

Benchmark paper · Leader — Claude Sonnet 4.5 (Claude Code) · 40.2 · Repository Generation · updated 2026-08-04

104
repository tasks
9
library categories
18.8k
average input tokens
40.2
top average test pass rate

Overview

NL2Repo-Bench starts where function completion and repository repair benchmarks stop. An agent receives one natural-language document and an otherwise empty workspace, then must choose an architecture, create every file, manage dependencies, implement cross-module behavior, and package an installable Python library. No source code, scaffold, or test case is visible during the main evaluation. The resulting repository is installed in a controlled image and checked against the original project's upstream pytest suite.

The 104 tasks cover nine categories: web development, testing, utility libraries, machine learning, data analysis and processing, database interaction, networking, batch file processing, and system tools. Input documents average approximately 18,800 tokens. Difficulty is based on original source size: 26 Easy repositories at no more than 1,500 LOC, 46 Medium at 1,500–4,000 LOC, and 32 Hard at 4,000 LOC or more. Candidate repositories had to be recent, have at least ten stars, and pass their own tests before inclusion.

Specifications are produced through reverse engineering rather than free-form summarization. An AST scanner inventories functions, classes, signatures, and locations; annotators describe the project, environment, file structure, API behavior, and critical implementation nodes; static checks compare the document with the inventory. Preliminary agent failures are then reviewed by a senior Python engineer so omissions in the document or container can be repaired instead of being counted as model failures.

Why it matters

Repository generation compounds errors across time and files. A plausible parser implementation can still break a grader, packaging can invalidate otherwise correct modules, and an early architectural choice can make later behavior impossible to reconcile. Averaging hidden test outcomes captures partial functional progress, while the complete-suite Pass@1 count shows whether that progress ever becomes a deliverable repository.

The benchmark also makes the information boundary explicit. When the paper reveals all hidden tests to Claude Sonnet 4.5 in Claude Code, average pass rate rises from 40.2% to 59.4% and fully passed repositories increase from 3 to 18. The remaining forty-point gap under that deliberately easier condition suggests that missing requirements are only part of the problem; coordination and large-scale implementation remain limiting.

Contributions

  • Formalizes from-scratch repository generation from a single requirements document: the agent receives no scaffolding, function signatures outside the document, source files, or visible tests.
  • Builds 104 quality-controlled Python-library tasks across nine categories, with specifications averaging 18,800 tokens and difficulty splits based on original repository size.
  • Pairs each task with a verified Docker image and the unmodified upstream pytest behavior, yielding execution-based evaluation rather than an LLM or qualitative judge.
  • Benchmarks twelve model-framework configurations and analyzes long-horizon failure modes including early stopping, non-finish behavior, tool-use allocation, context pressure, and iteration limits.

Method & evaluation

  • Candidate repositories must contain 300–120,000 lines of code, have at least 10 GitHub stars, have been created or updated within three years, include pytest tests, and pass their full native suite in the curated environment.
  • Annotators reverse-engineer each repository into a four-part document: Project Description, Supports, API Usage Guide, and Implementation Nodes. An AST scanner inventories classes, functions, signatures, and locations to check specification coverage.
  • Human experts review signatures and behavior, static checks compare documented APIs with the source inventory, and preliminary agent runs are inspected by a senior Python engineer to remove specification or environment faults.
  • At evaluation time the workspace contains only the task document. The agent may use its available tools without human follow-up or a fixed iteration cap, then submits a package that is installed inside the task's controlled test image.
  • The hidden upstream pytest suite executes all collectable cases even when some collection errors occur. Per-task test pass rate is averaged over 104 tasks; a separate Pass@1 count records tasks for which every test passes.

Evaluation metrics

Average Test Pass Rate — Higher is better. Mean percentage of hidden upstream pytest cases passed across the 104 generated repositories; this is the paper's Overall Score. Score range: [0, 100].

Repository Pass@1 — Higher is better. Count of repositories for which the generated package passes the complete upstream test suite in one run. Score range: [0, 104].

Difficulty Pass Rate — Higher is better. Average test pass rate split by original project size: Easy at no more than 1,500 LOC (26 tasks), Medium at 1,500–4,000 LOC (46), and Hard at at least 4,000 LOC (32). Score range: [0, 100].

Leaderboard

All rows use the same 104 specifications and hidden upstream tests. OpenHands is the default framework; Claude Code and Cursor variants are labeled explicitly. Gemini 3 Pro uses Cursor because the paper reports agent-loop errors in OpenHands.

#ModelAverage Test Pass Rate (%)
1Claude Sonnet 4.5 (Claude Code)40.2
2Claude Sonnet 4.5 (OpenHands)39.9
3Claude Sonnet 4.5 (Cursor)39.2
4Claude Sonnet 4 (OpenHands)37
5Gemini 3 Pro (Cursor)34.2
6DeepSeek-V3.2 (OpenHands)27.6
7Kimi-K2 (OpenHands)22.7
8DeepSeek-V3.1 (OpenHands)22.2
9GPT-5 (OpenHands)21.7
10Qwen3-235B-Instruct (OpenHands)17.9
11GLM-4.6 (OpenHands)17.5
12Qwen3-235B-Thinking (OpenHands)13.8

Reading the results

Claude Sonnet 4.5 occupies the top three rows across Claude Code (40.2), OpenHands (39.9), and Cursor (39.2). The one-point range supports the paper's claim that, for this model and setup, changing the harness matters less than changing the underlying model. Claude Sonnet 4 follows at 37.0 and Gemini 3 Pro in Cursor at 34.2; DeepSeek-V3.2 is the strongest remaining OpenHands system at 27.6.

Average test pass rate and complete-repository success tell different stories. The leaderboard winner fully passes only 3 of 104 repositories, while the lower-scoring Claude Sonnet 4 fully passes 5. For the winning system, average pass rate falls from 51.8% on Easy tasks to 44.5% on Medium and 25.1% on Hard. The benchmark therefore shows widespread partial implementation but very little reliable end-to-end completion, particularly once the original codebase exceeds 4,000 lines.

Figures & tables

Figure 1 (paper, PDF p. 2): A task specification showing the project overview, pinned environment, expected repository structure, API usage guide, and detailed implementation nodes supplied to the agent. Source: paper.
Figure 1 (paper, PDF p. 2): A task specification showing the project overview, pinned environment, expected repository structure, API usage guide, and detailed implementation nodes supplied to the agent. Source: paper.
Figure 2 (paper, PDF p. 5): Four-stage construction pipeline covering repository selection, reverse-engineered and AST-assisted document writing, test-image creation, and human/static/agent-based verification. Source: paper.
Figure 2 (paper, PDF p. 5): Four-stage construction pipeline covering repository selection, reverse-engineered and AST-assisted document writing, test-image creation, and human/static/agent-based verification. Source: paper.