Abstract
YuE is an open foundation-model family for long-form lyrics-to-song generation, not a benchmark. Its two-stage autoregressive design first predicts separate vocal and accompaniment tokens from lyrics, structural labels, style metadata, and optional reference audio, then fills residual codec codebooks before reconstructing 44.1 kHz audio. Segment-level conditioning refreshes lyric and section information throughout a song, while delayed in-context-learning data reduces direct copying from reference audio. The 7B Stage-1 model is trained to 1.75 trillion tokens and annealed for 40 billion more; the 1B Stage-2 model receives a 2-trillion-token budget. Training data includes 70,000 hours of English and Chinese speech and 650,000 hours of in-the-wild music, roughly 10% with retained lyrics. Human evaluation compares 42 English songs per system with four January 2025 commercial services.
Contributions
- Introduces track-decoupled next-token prediction, interleaving vocal and accompaniment tokens at every audio frame so each track remains explicit while both are modeled in one standard autoregressive sequence.
- Uses structural progressive conditioning: songs are segmented into about 14 sections on average, and each verse, chorus, bridge, intro, or outro refreshes its label, lyrics, and audio boundary tokens to limit long-context condition decay.
- Redesigns music in-context learning around 20–40-second single- or dual-track references without requiring a transcript, activates it only during annealing, and supports continuation-independent style, vocal, and accompaniment conditioning.
- Scales a two-stage LLaMA2-based model on large speech and music corpora, releases the training and inference implementation, and documents unsuccessful acoustic-token, unconditional-pretraining, and early-ICL choices rather than omitting them.
- Evaluates full-song musicality, control, acoustic distribution, vocal range, duration, alignment, multilingual adaptation, representation quality on MARBLE, and memorization against established datasets and commercial systems.
Method & evaluation
- X-Codec converts 16 kHz audio at 50 frames per second using 12 residual-VQ codebooks of size 1024; YuE uses the first eight. Stage 1 predicts semantic-rich codebook-0 tokens, Stage 2 predicts codebooks 1–7, and a lightweight Vocos-based module upsamples reconstruction to 44.1 kHz.
- Dual-NTP inserts a vocal token followed by an accompaniment token for each frame. About 40% of music tracks are separated with an ensemble of htdemucs_ft, Kim_Vocal_1, and UVR-MDX-NET-Inst_3, giving the model an explicit source-separation prior.
- All-in-one segmentation supplies intro, verse, chorus, bridge, and outro boundaries. The sequence interleaves each section label and its lyric text with start/end-of-audio tokens and dual-track audio codes; forced decoding inserts the next user-provided section prompt after every generated segment.
- A 20–40-second reference can contain vocals, accompaniment, a mixture, or both separated tracks. Reference codes are prepended only during the final annealing phase; earlier activation was rejected because the model copied the prompt and lost lyric control.
- Stage-1 training warms up over 280B tokens, reaches 1T tokens at constant learning rate, adds 750B tokens at 16K context, and anneals on 40B high-quality control tokens. The 7B scaling run uses as many as 512 H800 GPUs; Stage 2 is a 1B model trained with an 8K context and a 2T-token budget.
- The main evaluation generates 42 English full-length songs per system from shared genre, instrument, emotion, lyric, tempo, and 30-second chorus-reference prompts. Forty blind raters—including 12 speech/music-AI experts and seven trained musicians—make pairwise judgments against Suno V4, Udio, Hailuo, and Tiangong.
Evaluation metrics
Human pairwise preference — Higher is better. Blind A/B win, tie, and loss rates across overall musicality, vocal and accompaniment quality, arrangement, melody, structure, lyric adherence, genre, instrumentation, emotion, and tempo/rhythm control. Lyrics following is judged manually because Whisper was not reliable enough. Score range: [0, 100].
CLaMP 3 alignment — Higher is better. Music-specific text–audio representation similarity. YuE records 0.240 in Table 3; across systems, this metric correlates with human lyric, genre, instrumentation, emotion, and tempo-control preferences more consistently than CLAP.
Frechet Audio Distance — Lower is better. Distance between generated and reference audio-feature distributions using the paper's audioldm_eval setup. The authors caution that FAD has sample-size and domain biases and correlates weakly with their vocal/accompaniment quality judgments. Score range: [0, ∞).
Song-level vocal range — Higher is better. Pitch span in semitones estimated with RMVPE after filtering notes shorter than 40 ms and manually checking outputs. It measures vocal agility; YuE's median is about 27 semitones and the cross-system correlation with musicality is 0.857.
Figures & tables