Research · Environment Learning

EdgeBench

Day-scale environment learning measured through repeated interaction, hidden judging, and 134 real tasks.

Benchmark paper · Leader — Claude Opus 4.8 (Claude Code, 1M) · 51.3 · Environment Learning · updated 2026-08-04

134
real-world tasks
12h
main budget per trial
~38k
agent interaction hours
51.3
top Overall Score @12h

Overview

EdgeBench measures a different quantity from conventional endpoint benchmarks: how much an agent improves while working in an unfamiliar environment. Its 134 tasks span scientific research and machine learning, production software, open-ended optimization, professional deliverables, machine-checked theorem proving, and interactive games. Every task is designed to sustain at least 12 hours of work. The recorded human expert effort averages 57.2 hours per task and reaches 320 hours, indicating that the time budget reflects the task contract rather than an artificial delay.

The evaluation exposes two feedback loops. Inside the work container, an agent can run compilers, simulators, development splits, proof checkers, or game episodes as often as needed. Authoritative evaluation is separate: a submission is sent to an isolated judge with hidden tests, unseen seeds, private data, or client-style rubrics, and only the configured score or diagnostic feedback returns. Host-side snapshots score the trajectory at fixed intervals without revealing those measurements to the agent, so progress can be reconstructed even between explicit submissions.

The main experiment runs five frontier agent systems three times on every task for 12 hours, then reports best-so-far performance at two-hour intervals. The systems are not harness-identical: GPT models use Codex with a 256k compact window, GLM-5.1 and DeepSeek-V4-Pro use Claude Code with 200k, and Claude Opus 4.8 primarily uses Claude Code with a 1M window. The leaderboard therefore compares the paper's complete evaluated systems under the same tasks and duration, not isolated model APIs.

Why it matters

A final score cannot distinguish prior knowledge from adaptation. EdgeBench records the full trajectory, allowing researchers to ask whether an agent uses feedback to preserve gains, revise a strategy, and accumulate experience instead of merely sampling more attempts. The paper's restart ablation reinforces this distinction: continuous experience outperforms spending the same 12-hour budget on independent restarts across the tested slice.

The benchmark also makes long-horizon evaluation operationally concrete. Hidden assets are isolated from the work container, feedback is submission-gated, asynchronous judges can run while the agent continues working, and infrastructure incidents are tracked. Those choices matter at day scale, where service interruptions, context compaction, and evaluator leakage can otherwise be mistaken for model behavior.

Contributions

  • Introduces 134 day-scale executable tasks across six capability families, with 51 tasks publicly released and all tasks designed to sustain at least 12 hours of interaction.
  • Implements a dual-loop harness: agents iterate freely against local tools and visible development feedback, then submit artifacts to an isolated judge holding hidden tests, seeds, or grading criteria.
  • Collects three 12-hour trajectories for each task-model pair across five frontier agent systems, producing roughly 38,000 hours of interaction data and fixed-interval evaluator-only snapshots.
  • Shows that benchmark-average learning follows a three-parameter log-sigmoid law, remains stable on 28- and 72-hour subsets, and that frontier two-hour learning speed increased about eightfold over 221 days.

Method & evaluation

  • Experts selected unsaturated tasks that reward continued experimentation; the final taxonomy contains 39 science/ML, 36 systems/software, 19 optimization, 19 knowledge-work, 13 formal-proof, and 8 game tasks.
  • The agent works in a writable container with compilers, simulators, development splits, proof states, or game APIs. Hidden evaluation assets remain in a separate judge container mediated by a host-side server.
  • Agents may submit throughout a run and receive task-specific scores or diagnostics; the harness also evaluates snapshots at fixed intervals without exposing those snapshot scores to the agent.
  • Each of five systems receives three independent 12-hour runs per task. GPT-5.5 and GPT-5.4 use Codex with a 256k compact window; GLM-5.1 and DeepSeek-V4-Pro use Claude Code at 200k; Opus 4.8 primarily uses Claude Code at 1M.
  • For every elapsed-time budget, the benchmark takes each trajectory's best-so-far task score and averages across tasks. The paper fits S(t)=Smax/[1+(tmid/t)^β] to the resulting aggregate curves.

Evaluation metrics

Overall Score @12h — Higher is better. Benchmark-average best-so-far performance after 12 hours; task results are aggregated from up to three valid independent runs. Score range: [0, 100].

Category Score @12h — Higher is better. The same 12-hour score reported separately for Science, Code, Optimization, Knowledge, Math, and Games. Score range: [0, 100].

Log-sigmoid fit R² — Higher is better. Goodness of fit between the observed aggregate learning trajectory and the three-parameter log-sigmoid curve; all five 134-task fits are at least 0.997. Score range: [0, 1].

Leaderboard

All five rows use the 134-task, 12-hour setting from Table 2. The paper evaluates complete agent systems with different harnesses and context windows, so these numbers should not be read as a model-API-only comparison.

#ModelOverall Score @12h
1Claude Opus 4.8 (Claude Code, 1M)51.3
2GPT-5.5 (Codex, 256k)48.4
3GPT-5.4 (Codex, 256k)39.3
4GLM-5.1 (Claude Code, 200k)37.4
5DeepSeek-V4-Pro preview (Claude Code, 200k)31

Reading the results

Claude Opus 4.8 leads the common 12-hour table at 51.3, followed by GPT-5.5 at 48.4. GPT-5.4 reaches 39.3, GLM-5.1 37.4, and DeepSeek-V4-Pro preview 31.0. Every system improves materially after the two-hour mark: Opus rises from 39.0 to 51.3, GPT-5.5 from 36.8 to 48.4, GPT-5.4 from 29.7 to 39.3, GLM-5.1 from 26.0 to 37.4, and DeepSeek from 23.3 to 31.0. A two-hour leaderboard would therefore discard between 7.7 and 12.3 points of observed learning.

The aggregate curves are unusually regular despite jagged individual tasks. Across all 134 tasks, the five log-sigmoid fits have R² values of at least 0.997 and a mean of 0.998; fits on 28-hour and 72-hour subsets remain at least 0.993. That is evidence about population-level trajectories, not a claim that any single task improves smoothly. Category scores also differ sharply—Opus reaches 67.4 in Code but 36.5 in Optimization—so the overall score should be paired with the family breakdown when diagnosing a system.

Figures & tables

Figure 2 (paper, PDF p. 2): EdgeBench's 134 tasks split into six capability families—39 scientific/ML, 36 systems/software, 19 optimization, 19 professional knowledge-work, 13 formal-proof, and 8 game tasks—with representative task examples. Source: paper.
Figure 2 (paper, PDF p. 2): EdgeBench's 134 tasks split into six capability families—39 scientific/ML, 36 systems/software, 19 optimization, 19 professional knowledge-work, 13 formal-proof, and 8 game tasks—with representative task examples. Source: paper.
Figure 3 (paper, PDF p. 4): The continuous local loop and submission-gated judge loop, plus concrete local and hidden feedback channels for software, science, knowledge-work, and game tasks. Source: paper.
Figure 3 (paper, PDF p. 4): The continuous local loop and submission-gated judge loop, plus concrete local and hidden feedback channels for software, science, knowledge-work, and game tasks. Source: paper.