← overview

Song models — vocals only, with and without stem separation

Eight arms × 10 genres × 3 seeds, each judged twice: as generated, and as an isolated vocal stem. All audio 160 kbps stereo.

First: the setting that invalidated the original benchmark

dcw_enabled defaults to True in ACE-Step's GenerationParams — a wavelet-domain correction applied at every sampler step. Nobody set it. With it on, Whisper transcribes a 30-second XL-SFT clip as the single word "so". With it off, the same prompt and seed transcribe as the actual lyrics. Ruled out first, each by direct test: guidance scale (broken identically at cfg 1.0, where CFG is effectively off), step count, the 4B LM, shift, sampler, ADG, velocity clamping, checkpoint integrity. Only DCW moved it.
modelDCWprecisionrecallF1 what Whisper heard
XL SFT (4B)on — benchmarked0.5000.0110.021"so"
XL SFT (4B)off0.9880.8500.911"Oh, silent hour, descend on me…"
2B SFTon0.5000.0110.021"so"
2B SFToff0.9740.6660.790
XL turboon — benchmarked0.8540.4820.605"silent owl… lamps are loose"
XL turbooff0.9880.7000.818

It destroys the CFG models and degrades turbo. That is why the original 4.27/5 turbo score looked plausible and went unquestioned.

Stem separation — and what it costs in bulk

The prompt asks for one voice alone, and the models mostly comply — but "mostly" hides two different failures: bad singing, and good singing buried under a pad. Separating the vocal first lets the rubric grade the voice rather than the arrangement.

Three separators, deliberately: separation artefacts are model-specific, so agreement across architectures is what distinguishes a real effect from one model hallucinating.

separatorarchitectureclipsloadmean per 30 s clip× realtimeGPU-h per 10k clipsGPU-h per 1 M clips
BS-Roformerbs_roformer24011 s1.55 s19.4×4.3430
MelBand Roformermel_roformer24010 s1.13 s26.6×3.1314
Demucs v4 fthtdemucs_ft2402 s5.43 s5.5×15.11,508

Measured on one GH200, 240 clips each. The Roformers are 4.8× faster than Demucs here (27× realtime against 6×). At 1.13 s per clip, separating a million 30-second clips costs about 314 GPU-hours — roughly 0.1 % of the 226,935 GPU-h the full 6,000-voice speech corpus would cost. Separation is cheap relative to generation; it is not free.

The same rubric, on the mix and on the isolated vocal

Read score B carefully — it does not mean the same thing in both columns.
A — singing quality & prompt match. Meaningful on both. On the stem the judge hears the voice without accompaniment masking or flattering it. This is the comparison worth making.
B — is it one person alone? On the mix this measures the model: did it obey "no instruments, no second voice". On the stem it can no longer measure instruments at all — the separator removed them by construction. It still measures extra voices, since a vocal separator pulls voice away from instruments, not voice away from voice.
So B(mix) → B(stem) decomposes the failure: what B recovers after separation was instrumental bleed; what stays low is a genuine second singer. B(stem) rising is not the model getting better.
armsettingnA mixA vocalsΔAB mixB vocalsΔB
ACE-Step XL SFTDCW on — the broken default300.370.33-0.030.600.23-0.37
ACE-Step XL SFTDCW off — fixed304.904.87-0.032.974.17+1.20
ACE-Step XL turboDCW on304.203.80-0.402.333.63+1.30
ACE-Step XL turboDCW off — fixed304.374.27-0.102.673.80+1.13
ACE-Step 2B SFTDCW off304.804.80+0.003.734.50+0.77
ACE-Step 2B turboDCW off304.074.57+0.502.734.40+1.67
HeartMuLaunchanged303.073.90+0.832.204.60+2.40
LeVo 2unchanged304.404.33-0.074.074.07+0.00

Do the three separators agree?

If a score shift were a separation artefact it would differ by architecture. BS-Roformer covers every clip; the other two cover seed 0 as a cross-check, so their n is smaller and the comparison is indicative rather than matched.

audio judgednmean Amean B
the mix, as generated2403.772.66
vocal stem — BS-Roformer2403.863.67
vocal stem — MelBand Roformer803.923.55
vocal stem — Demucs v4 ft803.983.79

Listen — mix and isolated vocal, side by side

Seed 0 of each genre. Left is what the model produced; right is the BS-Roformer vocal stem of that exact take.

ACE-Step XL SFT — DCW on — the broken default

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

ACE-Step XL SFT — DCW off — fixed

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

ACE-Step XL turbo — DCW on

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

ACE-Step XL turbo — DCW off — fixed

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

ACE-Step 2B SFT — DCW off

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

ACE-Step 2B turbo — DCW off

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

HeartMuLa — unchanged

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

LeVo 2 — unchanged

genreas generatedvocal stem
Fragile intimate soul / jazz
Operatic / classical solo voice
Gangsta rap, solo
Pure a cappella
Smoky rock ballad, husky voice
Punk rock, shouted and raw
Gospel, full-voiced and spiritual
Folk / singer-songwriter, plain
Metal, harsh or belted
Children's song / lullaby

Method and limits


Separators: BS-Roformer (Viperx ep.317, SDR 12.98) · MelBand Roformer InstVoc Duality V2 · Demucs v4 htdemucs_ft.