Workflow-GYM
Specialized desktop work scored at the final artifact or GUI state after 30 to 110 dependent actions.
Overview
Workflow-GYM evaluates whether a computer-use agent can complete a professional objective inside specialized desktop software, not merely execute a short click sequence. Its 338 tasks cover data analysis, engineering and design, finance and management, geography and environment, multimedia and creative work, and scientific computing. These expand into 23 subdomains and 56 pinned virtual-machine environments, including CAD, GIS, medical imaging, circuit design, statistics, databases, audio/video production, and scientific simulation.
Seventy-one domain experts contributed workflows from their own practice. From more than 1,000 candidates, tasks were retained only if they required specialized knowledge and software, were not directly recoverable from a public tutorial, involved at least 30 atomic GUI actions, and had a closed success condition. The released set contains 129 Easy tasks with 30–44 expert steps, 159 Medium tasks with 45–60, and 50 Hard tasks with 61–110.
At test time the agent sees one natural-language objective and no intermediate hints. A hidden expert procedure proves that the task can be completed from the initialized VM, but the agent must discover its own path within 400 screenshot-action rounds. Evaluation checks either the produced artifact or final GUI state against task-specific criteria. Each model runs every task three times, so Avg Pass is calculated over 1,014 trials and Pass@3 records whether at least one attempt succeeds.
Why it matters
Professional GUI work is brittle in a way that atomic computer-use tasks are not. Choosing the wrong base object in a CAD workflow, omitting a layout stage, or drifting toward an intermediate subgoal can leave dozens of later actions locally plausible while making the final artifact irrecoverably wrong. Final-state verification captures that compounding behavior rather than rewarding isolated correct clicks.
The benchmark also separates planning from execution through guidance ablations. Expert step-by-step text improves every tested model, and video demonstrations add further gains on a controlled 100-task subset. Yet failures persist through localization errors, skipped instructions, repeated-action loops, and inability to recover once the live interface diverges from demonstrated frames. This makes Workflow-GYM useful for diagnosing both high-level workflow control and low-level GUI reliability.
Contributions
- Introduces 338 authentic professional GUI tasks across six domains, 23 subdomains, and 56 software-specific virtual machines, selected from more than 1,000 expert proposals.
- Targets genuinely long workflows: every task requires at least 30 atomic actions, with 129 Easy tasks at 30–44 steps, 159 Medium at 45–60, and 50 Hard at 61–110.
- Defines artifact- and GUI-state-based success criteria, hidden expert execution procedures, and a three-stage validation process covering environment, instruction, and end-to-end agent failures.
- Runs six frontier model-agent systems for three trials on all 338 tasks and analyzes outcome failures, stage omission, error propagation, objective drift, software-knowledge gaps, action loops, and procedural-guidance ablations.
Method & evaluation
- Domain experts propose workflows from their own professional practice. Tasks must require specialized software and domain knowledge, take at least 30 atomic GUI actions, resist retrieval from public tutorials, and admit an objective success criterion.
- Fifty-six full virtual-machine images provide pinned professional software. At runtime, task-specific inputs and configuration are injected, the required application is pre-launched, and unrelated login or setup steps are removed.
- Each task includes a self-contained objective, a final artifact or target GUI state, deterministic evaluation criteria, and a hidden atomic expert procedure. Experts replay that procedure to verify both environment and task solvability.
- At test time an agent sees only the objective and interacts through its model-specific GUI framework for at most 400 screenshot-action rounds. Every task is run independently three times for each system.
- Artifacts are checked with task-appropriate rules or LLM-based rubrics; final GUI states are judged from screenshots and validated criteria, using Seed-1.8 for non-rule-based cases. Avg Pass averages all 1,014 binary trial outcomes.
Evaluation metrics
Avg Pass — Higher is better. Percentage of successful trials across 338 tasks and three independent runs per task (1,014 binary outcomes). Score range: [0, 100].
Pass@3 — Higher is better. Percentage of tasks solved in at least one of the three independent attempts. Score range: [0, 100].
Difficulty Avg Pass — Higher is better. Average trial success split by hidden expert procedure length: Easy 30–44 steps, Medium 45–60, and Hard 61–110. Score range: [0, 100].
Leaderboard
All six rows use the updated 338-task, three-trial, 400-interaction setting from Table 2. Each model uses its aligned GUI-agent framework, so the results represent complete model-framework systems.
| # | Model | Avg Pass (%) |
|---|---|---|
| 1 | Gemini 3.1 Pro | 30.67 |
| 2 | Kimi K2.6 | 29.68 |
| 3 | Seed-2.0-Lite | 18.24 |
| 4 | GPT-5.4 | 17.85 |
| 5 | GPT-5.4-mini | 15.98 |
| 6 | Gemini 3 Flash | 7.89 |
Reading the results
Gemini 3.1 Pro leads Avg Pass at 30.67%, only 0.99 points ahead of Kimi K2.6 at 29.68%. The order reverses on Pass@3: Kimi reaches 41.42% and Gemini 41.12%, indicating that Gemini is marginally more consistent per trial while Kimi solves a slightly larger set at least once. Seed-2.0-Lite is third at 18.24%, ahead of GPT-5.4 at 17.85, GPT-5.4-mini at 15.98, and Gemini 3 Flash at 7.89.
Longer procedures reliably lower success. Gemini 3.1 Pro drops from 39.28% on Easy to 26.62% on Medium and 21.33% on Hard; Gemini 3 Flash falls from 12.14% to 6.08% and 2.67%. Domain results are also uneven: Kimi K2.6 reaches 45.24% on Data Analysis but 21.48% on Engineering and Design, while Gemini 3.1 Pro scores 39.88% and 24.44% respectively. No system solves even one third of all trials, so the leaderboard is measuring substantial unresolved headroom rather than near-saturation.
Figures & tables
Read the research note
Workflow-GYM: When One Missed Stage Invalidates the Job