Research · Foundation Model Architecture

Ouro

Model paper · Foundation Model Architecture · updated 2026-08-04

Abstract

Ouro is a family of 1.4B- and 2.6B-parameter foundation models built with a looped decoder architecture, not an evaluation benchmark. A shared stack of Transformer layers is applied repeatedly so recurrent depth adds latent computation without adding parameters or output tokens. An exit gate estimates a distribution over depths, and entropy regularization prevents training from collapsing to always using the deepest step; a second gate-only stage learns to stop when an additional loop no longer improves token loss. The base models are trained on 7.7 trillion tokens through stable pretraining, high-quality annealing, long-context training, and mid-training, then Thinking variants receive supervised reasoning tuning on 8.3 million examples. Evaluations use existing knowledge, reasoning, math, coding, safety, and trace-faithfulness tasks, with controlled synthetic studies separating factual storage from knowledge manipulation.

Contributions

  • Scales parameter-shared Looped Language Models to 7.7T-token foundation-model training and releases 1.4B and 2.6B base and Thinking variants trained for four recurrent steps.
  • Introduces an entropy-regularized expected language-modeling objective with a depth-unbiased uniform prior, followed by gate-only training that labels continuation from the measured token-loss improvement at each loop.
  • Documents an end-to-end open-data recipe: 6T-token stable pretraining, 1.4T-token continual-training annealing, 20B-token 64K-context training, 300B-token mid-training, and an 8.3M-example reasoning SFT stage.
  • Shows through controlled synthetic biographies that looping leaves knowledge storage near two bits per parameter but improves composition and multi-hop manipulation; separate experiments study recurrent-depth scaling, early exit, KV-cache reuse, safety, and reasoning-trace faithfulness.
  • Reports that the 2.6B R4 base model reaches 55.73 MMLU-Pro, 80.46 BBH, and 90.85 MATH500, while the 1.4B R4 model reaches 71.02 BBH, 78.92 GSM8K, and 82.40 MATH500 under the paper's shared evaluation harness.

Method & evaluation

  • Each decoder model uses multi-head attention with RoPE, SwiGLU feed-forward layers, sandwich RMSNorm, and a 49,152-token vocabulary. Ouro-1.4B has 24 layers at hidden size 2048; Ouro-2.6B duplicates the stack to 48 layers while preserving the same width.
  • The depth-L stack is reused for up to four recurrent steps at inference. A language-model head computes next-token loss after every step, while an exit gate converts the current hidden state into a per-step halting probability and supports threshold-based early exit.
  • Stage-I training minimizes the exit-probability-weighted language-model loss minus an entropy term, equivalent to a KL penalty toward a uniform prior over depths. Stage II freezes the language model and tunes only the gate from observed loss improvements to penalize both underthinking and overthinking.
  • The 7.7T-token base recipe uses 4K sequences for two 3T-token stable phases, 16K for 1.4T-token annealing, 64K for 20B long-context tokens, and 32K for 300B mid-training tokens. Training starts with eight loops, then moves to four for stability; batch size grows from 4M to 8M tokens.
  • Thinking models are supervised-tuned for two epochs at 32K context on 8.3M public examples: 3.5M mathematics, 3.2M code, 808K science, and 767K chat examples. Reported RLVR attempts did not improve on this SFT checkpoint and are not presented as part of the final recipe.
  • Base-model comparisons use lm-eval-harness and EvalPlus on shared prompts across MMLU, MMLU-Pro, BBH, ARC, commonsense, math, and code tasks. Thinking-model tests use one in-house harness, temperature 1.0, top-p 0.7, and a fixed judging rubric on AIME, OlympiadBench, BeyondAIME, HLE, GPQA, and SuperGPQA.

Evaluation metrics

Task accuracy — Higher is better. Percentage correct on multiple-choice and exact-answer suites such as MMLU, MMLU-Pro, BBH, ARC, GSM8K, MATH500, OlympiadBench, GPQA, and SuperGPQA under the paper's task-specific prompting settings. Score range: [0, 100].

AIME pass@1 / pass@10 — Higher is better. Success on the 30 integer-answer questions in each AIME 2024 and AIME 2025 set, reported for one sample and at least one success among ten samples. Thinking-model decoding uses temperature 1.0 and top-p 0.7. Score range: [0, 100].

EvalPlus pass@1 — Higher is better. First-sample execution success on HumanEval, MBPP, and their stricter HumanEval+ and MBPP+ test suites, evaluated with the same pipeline used for the compared base models. Score range: [0, 100].

Accuracy at exit depth — Higher is better. Task accuracy paired with average recurrent exit round or a fixed depth. It measures the compute–quality trade-off of static exit, hidden-state thresholds, the pretrained gate, and the separately trained adaptive gate rather than collapsing them into a single efficiency score. Score range: [0, 100].

Figures & tables

Figure 3 (paper, PDF p. 4): LoopLM training repeatedly applies a shared layer stack, computes language-model loss and exit probability at each recurrent step, and uses the learned cumulative exit probability for early termination at inference. Source: paper.
Figure 3 (paper, PDF p. 4): LoopLM training repeatedly applies a shared layer stack, computes language-model loss and exit probability at each recurrent step, and uses the learned cumulative exit probability for early termination at inference. Source: paper.
Table 8 (paper, PDF p. 14): Ouro-2.6B R4 compared with 3B–12B dense base models on general, mathematics, and coding tasks, including model size and reported pretraining-token count for each system. Source: paper.
Table 8 (paper, PDF p. 14): Ouro-2.6B R4 compared with 3B–12B dense base models on general, mathematics, and coding tasks, including model size and reported pretraining-token count for each system. Source: paper.