YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality
Writing a score before rendering audio turns out to matter. YuE2's central claim is that making melody, harmony, rhythm, and form explicit in a readable intermediate representation—before any acoustic generation begins—measurably improves what listeners actually hear. The same model checkpoint that composes an ABC-notation lead sheet then expands it through semantic music tokens and continuous acoustic latents into a finished stereo recording. Experts who compared planned versus unplanned generation from the same checkpoint, on the same prompts, preferred the planned output for overall quality by roughly 49% to 35%, and for musicality by 45% to 29%.
The architecture is a 28-layer Mixture-of-Transformers with about 3.58 billion parameters, jointly handling autoregressive symbolic and semantic prediction alongside bidirectional flow matching for acoustics. Training required two new components: MERT2, a music encoder that synthesizes discrete targets from two frozen encoder views before any fine-tuning begins, and SheetSage2, a full-song transcriber that recovers lead sheets from recordings to supply the symbolic training targets. MERT2 leads on 14 of 15 MARBLE metrics against all evaluated external baselines. SheetSage2-AR, trained entirely on labels its own prober generated, leads on 12 of 15 benchmark-metric pairs in full-song transcription—and surpasses its label generator on 10 of those 15 pairs, which is a notable result about autoregressive transcription outrunning its own supervision source.
On WildSongBench, evaluated across 192 prompts spanning 15 genre categories, YuE2 leads all evaluated public systems on SongBench Global Average under the standard two-candidate protocol. With eight candidates, it achieves the highest observed mean across every evaluated system, public or proprietary. Audio quality is where it most consistently beats commercial generators: experts preferred its audio quality over each of six proprietary systems, with a tie-adjusted average preference of 58.9% across those comparisons. Overall quality preferences are more mixed—Suno v6 receives more overall preferences than the best-of-8 setting, even though YuE2 scores higher on SongBench.
The score is not just an internal planning device. Because it is human-readable ABC notation, an external language model can revise it directly. The paper demonstrates this through a case study where a Mandarin pop song is iteratively reworked into English jazz through fourteen rendered versions, with a language-model agent translating user feedback into explicit score revisions. Score edits are also measurably local: harmony edits achieve roughly 80% target-chord agreement while retaining over 93% of the unedited vocal melody. Cover generation without any cover-specific training—just supplying a SheetSage2 transcription as a fixed score prefix—outperforms both evaluated dedicated cover systems on all eight work-identity retrieval measures.
The honest limitation is that relaxing score conditioning raises musicality and style-alignment scores while reducing work-identity retrieval. Harmony, it turns out, is what makes a cover sound like the original work, and also what constrains how far the style can shift. That tension is not resolved; it is documented.
Composing an explicit, editable score before rendering audio demonstrably improves what expert listeners hear—and makes the composition itself available for revision.
Sources & links
Related on SkillFed
A GNN-guided offline simulation pipeline pre-builds and validates Adobe Illustrator scripting skills, beating live gpt-4o code generation 44.7% to 28.7% on success rate while…
A capability-tree-plus-DAG framework holds top rank across 200-to-200,000-skill ecosystems, while handing an agent the same full skill pool unstructured gets worse, not better, as…
audio-router intelligently directs audio processing requests to the right specialized skill, whether for playback control, sound analysis, or reactive audio handling. Built for…