Abstract
AutoMV is a training-free multi-agent system for turning a complete song into a shot-based music video; the paper is principally a generation-system study, with a 30-song evaluation suite used to test it. Music preprocessing extracts caption, section structure, a separated vocal stem, and timestamped lyrics. A Gemini screenwriter converts that context into 3–15-second narrative units, a Doubao director specifies shots and keyframes, and a shared character bank carries appearance constraints across clips. The renderer routes general scenes to Doubao video generation and salient singing scenes to Wan-2.2 lip-sync generation. Gemini 2.5 Pro then filters up to three image or video candidates for physical plausibility, prompt alignment, and continuity. Evaluation covers professionally released English, Chinese, Japanese, and Korean songs, two closed commercial systems, and professional human-directed videos.
Contributions
- Defines a training-free, full-song production pipeline that joins music-information retrieval, script and shot planning, image/video generation, lip synchronization, verification, and final compilation instead of generating disconnected short clips.
- Introduces music-aware controls for long-form consistency: SongFormer section boundaries, htdemucs vocal separation, Whisper lyrics and timestamps, 1/24-second timing quantization, and a persistent character bank used by downstream prompts.
- Builds a system evaluation on 30 professionally released songs spanning English, Chinese, Japanese, and Korean, with 12 criteria grouped into Technical, Post-Production, Music Content, and Art and rated by industry experts and multimodal judges.
- Reports component ablations: removing lyric/timestamp context reduces content and audio–visual alignment; removing the character bank lowers character-consistency score from 3.07 to 1.22; removing the verifier reduces visual quality from 3.28 to 2.61.
Method & evaluation
- Qwen2.5-Omni captions genre, mood, instrumentation, and vocalist attributes; SongFormer segments musical structure; htdemucs separates vocals from accompaniment; Whisper transcribes the vocal stem, after which Gemini refines lyrics and timestamps.
- The Gemini screenwriter divides the song at lyric and section boundaries into user-configurable 3–15-second shots. Boundaries are quantized to 1/24 second so clip durations remain aligned when the full video is assembled.
- A Doubao director writes shot-level action, setting, and camera instructions and generates keyframes. Character descriptions covering appearance, age, gender, and attire are retrieved from a shared bank and injected into every relevant prompt.
- General cinematic shots use Doubao video generation; shots where mouth articulation matters use Wan-2.2 with the separated vocal stem for lip synchronization. Longer segments are composed from 3–8-second subclips, optionally reusing the preceding end frame.
- Gemini 2.5 Pro evaluates as many as three candidates per keyframe or clip. It rejects physically implausible outputs, then selects among surviving clips using instruction alignment and identity continuity before full-song compilation.
- The paper compares AutoMV, Revid.ai-base, OpenArt-story, and professional videos on the same song inputs. Experts watch each full video with audio; the automated study uses ImageBind plus full-video Gemini scoring under the same 12-item rubric.
Evaluation metrics
Expert weighted score — Higher is better. Industry practitioners rate 12 sub-criteria after watching each full music video with audio. Technical and Post-Production each contribute 20% of the total, while Music Content and Art each contribute 30%. AutoMV scores 2.42 versus 1.45 for OpenArt-story and 1.06 for Revid.ai-base. Score range: [1, 5].
Multimodal-judge score — Higher is better. Gemini receives the complete video and song and scores the same 12 criteria. Category scores are sub-criterion means and use the same 20/20/30/30 weighting; the paper separately reports correlation with human ratings rather than treating the judge as ground truth. Score range: [1, 5].
ImageBind audio–visual score — Higher is better. Embedding similarity between sampled video frames and the corresponding audio, reported as a percentage. AutoMV records 24.4, compared with 19.9 for Revid.ai-base, 18.5 for OpenArt-story, and 24.1 for professional videos. Score range: [0, 100].
Figures & tables