Abstract
OpenCodeInterpreter is a family of open code-generation systems trained to revise programs after compiler output and user feedback, rather than a new benchmark. Its Code-Feedback dataset contains 68K multi-turn dialogues and 192K turns assembled through query packing, simulated interactions, deliberate code correction, and related or follow-up LeetCode problems. Models based on CodeLlama, DeepSeekCoder, and StarCoder2 are instruction-tuned on this feedback data mixed with high-quality single-turn examples. Evaluation uses HumanEval, MBPP, and their EvalPlus extensions in single-turn and two-round refinement settings. OpenCodeInterpreter-DeepSeekCoder-33B scores 79.0 average pass@1 (70.4 on the Plus suites) before feedback and 83.2 (76.4) with execution feedback; synthesized human feedback raises it to 88.0 (81.0), while the oracle-feedback condition reaches 91.6 (84.6).
Contributions
- Releases Code-Feedback with 68K dialogues and 192K turns, explicitly representing multi-turn conversation, execution feedback, and human feedback rather than only static instruction-answer pairs.
- Builds five data streams: 16K packed dialogues, 51K simulated interactions, 0.5K code-correction dialogues, 0.3K similar-problem dialogues, and 0.2K LeetCode follow-up dialogues.
- Trains open interpreter variants across CodeLlama, DeepSeekCoder, and StarCoder2 scales so the same model can generate code, consume execution diagnostics, and revise its answer within the conversation.
- Reports controlled feedback gains: the 33B DeepSeekCoder variant improves from 79.0 to 83.2 average pass@1 with execution feedback and to 88.0 with synthesized human feedback; the oracle-feedback result is 91.6.
- Tests real multi-turn behavior on the 10 coding prompts in MT-Bench, where OpenCodeInterpreter-DS-33B scores 6.8 overall versus 5.5 for DeepSeekCoder-33B-Instruct under GPT-4 judging.
Method & evaluation
- The authors aggregate 287K queries from four open code-instruction datasets, score complexity twice with Qwen-72B-Chat, and retain 156K queries rated 4 or 5 as a challenging seed pool.
- Single-turn packing groups semantically related queries; GPT-3.5/GPT-4 interaction simulation executes generated code and adds one of ten user-feedback categories; deliberate code-correction traces teach recovery from diagnostics; LeetCode tasks supply related and follow-up problems.
- Base models are fine-tuned for three epochs with learning rate 2e-5, 5% warm-up, cosine scheduling, and a 4096-token cutoff. Code-Feedback is mixed with WizardCoder-110K single-turn data at the selected 1:2 feedback-to-single-turn ratio.
- At inference, generated code is sanitized and executed. Failed solutions receive exception text, mismatched expected/actual outputs, or a timeout message; evaluation stops on success or after two refinement rounds.
- HumanEval, MBPP, HumanEval+, and MBPP+ are decoded greedily and scored with EvalPlus pass@1. Multi-turn conditions isolate execution feedback, GPT-4-synthesized user feedback, and an oracle variant whose feedback generator can see the reference solution; MT-Bench adds 10 two-turn coding conversations.
Evaluation metrics
EvalPlus pass@1 — Higher is better. Percentage of HumanEval, MBPP, HumanEval+, or MBPP+ tasks solved by the single greedily decoded candidate; refinement settings allow at most two feedback rounds. Score range: [0, 100].
Average pass@1 — Higher is better. Arithmetic mean of HumanEval and MBPP pass@1. The parenthesized Average+ reported by the paper is the corresponding mean of HumanEval+ and MBPP+. Score range: [0, 100].
MT-Bench coding score — Higher is better. GPT-4 judgment score on 10 coding conversations, reported separately for the first and follow-up turns and as their average. Score range: [1, 10].
Figures & tables