MiniMax Music 3 vs LeVo 2 — one voice, no instruments
Twenty prompts spread across the measured genre space, the same prompt to both models, vocals only. Every clip judged blind, twice, by Gemini 3.6 Flash on five anchored axes, and the "no instruments" claim checked against BS-Roformer stem separation rather than left to the judge. All audio 160 kbps stereo.
0.45re-judging the same clip moves overall by this much
0.3751MiniMax — accompaniment fraction
0.0000LeVo vocal — accompaniment fraction
Verdict.LeVo 2 wins overall: 1.62 vs 3.20, a paired difference of 1.57 ± 0.35 against a re-judging spread of 0.45 on the same axis, winning 18 of the 20 prompts, so the difference survives the measurement's own noise. Axis by axis, the differences that clear their own noise floor are: singing quality 2.90 vs 3.77, paired difference -0.88 ± 0.18 (noise floor 0.45; 2 prompts to 15, 3 tied); no instruments 2.55 vs 4.65, paired difference -2.10 ± 0.23 (noise floor 0.23; 0 prompts to 18, 2 tied); genre match 2.27 vs 3.17, paired difference -0.90 ± 0.37 (noise floor 0.55; 3 prompts to 16, 1 tied); overall 1.62 vs 3.20, paired difference -1.57 ± 0.35 (noise floor 0.45; 2 prompts to 18, 0 tied). Not separable: one voice alone. Objectively, BS-Roformer assigns 0.3751 of MiniMax's separated energy to something other than the voice, against 0.0000 for LeVo 2 in its a cappella mode.
Read the LeVo column with this in mind. LeVo 2's a cappella clips are rendered with gen_type='vocal', which never decodes the model's bgm codebook at all: an accompaniment is not suppressed by the prompt, it is structurally impossible. MiniMax Music 3 has no equivalent switch, so for it "no instruments" is a request the model can refuse. The third column is the control — the identical prompt given to LeVo 2 with gen_type='mixed', where it must comply by prompt alone, and where its accompaniment fraction rises to 0.5812. Comparing MiniMax against that column is the like-for-like prompt-compliance test; comparing it against the vocal column is a comparison of what each system can be made to deliver.
Against that control, the picture changes. Paired, MiniMax minus LeVo 2 in mixed mode: singing quality -0.82 ± 0.18 (noise 0.45); one voice alone +0.25 ± 0.46 (noise 0.58); no instruments +0.55 ± 0.22 (noise 0.23); genre match -0.10 ± 0.35 (noise 0.55); overall +0.33 ± 0.24 (noise 0.45). The differences that clear both their noise floor and twice their standard error are singing quality, no instruments. So when both models have to keep the instruments out by prompt alone, MiniMax is the better of the two at doing it — it just sings less well. Neither is close to a cappella in that condition.
Per-axis scores
Five axes, each anchored in words at every level rather than left as a bare 0-5 scale. S is the singing alone; V counts voices and ignores instruments; I counts non-vocal sound and ignores how many singers there are; G is fit to the requested style, age and gender; O is one holistic judgement. V and I are deliberately separated: the older two-axis rubric folded them together, which is exactly the confound the stem-separation row below exists to break. Each arm appears twice — as generated, and as its BS-Roformer vocal stem.
arm
n
singing quality S
one voice alone V
no instruments I
genre match G
overall O
MiniMax Music 3
20
2.90 ±0.68
3.73 ±1.22
2.55 ±0.93
2.27 ±1.35
1.62 ±0.97
— vocal stem
20
2.33 ±0.92
4.15 ±1.27
4.45 ±0.94
2.10 ±1.17
1.93 ±1.04
LeVo 2 (gen_type=vocal)
20
3.77 ±0.47
4.28 ±1.15
4.65 ±0.76
3.17 ±1.68
3.20 ±1.37
— vocal stem
20
3.77 ±0.66
4.28 ±1.11
4.95 ±0.22
3.25 ±1.54
3.20 ±1.33
LeVo 2 (gen_type=mixed)
20
3.73 ±0.44
3.48 ±1.26
2.00 ±0.16
2.38 ±1.44
1.30 ±0.41
— vocal stem
20
3.55 ±0.74
3.70 ±1.45
4.97 ±0.11
3.23 ±1.42
2.95 ±1.23
Mean over n clips ± standard deviation across clips. Each clip's score is the mean of two independent judging passes. Green ≥ 4.00, red < 3.00.
Paired, prompt by prompt
Both arms saw byte-identical prompts on the same 20 genres, so the comparison can be paired: for each prompt, MiniMax minus LeVo. Pairing removes the between-genre variance, which is much larger than the between-model variance and would otherwise bury it. A positive number favours MiniMax.
comparison
axis
mean difference
± s.e.
prompts won by MiniMax
tied
won by LeVo
MiniMax − LeVo vocal (n=20)
singing quality
-0.88
0.18
2
3
15
one voice alone
-0.55
0.42
4
6
10
no instruments
-2.10
0.23
0
2
18
genre match
-0.90
0.37
3
1
16
overall
-1.57
0.35
2
0
18
MiniMax − LeVo mixed (n=20)
singing quality
-0.82
0.18
1
5
14
one voice alone
+0.25
0.46
10
4
6
no instruments
+0.55
0.22
6
13
1
genre match
-0.10
0.35
7
4
9
overall
+0.33
0.24
6
11
3
How much of this is noise
Every clip was judged twice, in two independent passes over a shuffled queue, 120 clips in both. The table is the disagreement between those two passes: it is the smallest difference this measurement can see. Any gap in the table above that is narrower than the corresponding row here is noise, and the verdict treats it as such.
axis
n clips
mean |pass1 − pass2|
identical
within 1 point
r
singing quality S
120
0.45
61%
95%
0.668
one voice alone V
120
0.58
67%
82%
0.682
no instruments I
120
0.23
87%
93%
0.882
genre match G
120
0.55
54%
92%
0.847
overall O
120
0.45
66%
92%
0.823
The objective check: accompaniment fraction
The judge is a language model listening to audio; it can be wrong about whether a guitar is there. acc_frac is not: it is BS-Roformer's own energy split, ||instrumental||² / (||vocal||² + ||instrumental||²), the share of separated energy the separator assigned to something that is not a voice. A bare voice sits near zero; an ordinary pop mix is 0.5–0.9. The last column is the diagnostic that matters: if V (one voice alone) does not improve once the instruments are stripped out, whatever the judge heard was a second singer, because a vocal separator pulls voice away from instruments and never voice away from voice.
arm
n
mean
median
min
max
clips > 0.05
judge I (mix)
judge V mix → stem
MiniMax Music 3
20
0.3751
0.4028
0.0578
0.7159
20
3.73
3.73 → 4.15 (+0.43)
LeVo 2 (gen_type=vocal)
20
0.0000
0.0000
0.0000
0.0002
0
4.28
4.28 → 4.28 (+0.00)
LeVo 2 (gen_type=mixed)
20
0.5812
0.5847
0.3189
0.7879
20
3.48
3.48 → 3.70 (+0.23)
Method
How the 20 prompts were sampled
The frame is the 153-genre solo-voice taxonomy built earlier in this project, reduced to
50 clusters by Ward linkage on VoiceCLAP genre centroids; each cluster is represented by
its medoid, a real genre rather than a synthetic average. From those 50 medoids the 20
here were chosen by farthest-point sampling on the same cosine geometry — seed the
largest cluster, then repeatedly take the medoid whose minimum distance to everything already
chosen is largest. No RNG, so the set is reproducible from
song/scripts/mm_prompts.py.
The result is a spread rather than 20 variations of one thing: minimum pairwise cosine
distance inside the 20 is 0.303, against 0.082 across all 50 medoids, so no two
prompts are near-duplicates. It covers 10 of the 11 families present among the medoids,
all five age bands (one child, two teen, four young adult, twelve adult, one older) and all
three gender targets. Family and age coverage were checked, never imposed: they
fall out of the geometry.
One prompt, given to both
Every genre has a single prompt string handed verbatim to both models, and one lyric shared
across all 20 genres so that any measured difference is the singing and not the words.
A ceiling set by LeVo, applied to both. LeVo 2 feeds the description to its
type_info conditioner, configured max_len: 100 and truncating with
x[:, :max_len] — a silent tail cut. One token is the <|im_start|>
it prepends, leaving 99. This project's standard vocals-only caption tokenises to 94–109
Qwen tokens, so 9 of these 20 prompts would have been cut, several of them losing the
words “no choir or background vocals”. The caption was therefore rewritten compactly
to 57–65 tokens, keeping every element (solo, a cappella, gender, age, no
instruments, no backing track, no second voice, no harmony, no choir, the genre clause, dry
close-mic) — and the same short caption went to both models. Shortening is prompt
tailoring, so it was applied to both arms or it would not have been applied at all.
The lyric syntax differs, and has to. The words are identical; only the separators
change, because the two checkpoints have incompatible input contracts:
MiniMax Music 3 — each [tag] on its own line. Text sharing a line
with a leading tag is dropped by the checkpoint
(diffusers/modular_pipelines/minimax_music3/encoders.py::_normalize_lyrics).
LeVo 2 — one line, ' ; ' between sections, '.' between
phrases.
The one lyric, in both forms, in full:
MiniMax Music 3
[verse]
I left the window open half the night
counting all the reasons I should stay
the morning came in quiet as a hand
and took the argument away
[chorus]
So carry me, carry me home
over the water and back through the town
carry me, carry me home
before the light goes down
LeVo 2
[verse] I left the window open half the night. counting all the reasons I should stay. the morning came in quiet as a hand. and took the argument away. ; [chorus] So carry me, carry me home. over the water and back through the town. carry me, carry me home. before the light goes down.
What was reused and what was generated
Nothing was reused. All 60 clips were generated for this comparison.
The existing LeVo 2 corpora were checked rather than assumed: all 20 genres appear in
TTS-AGI/levo2-vocals-captioned (8,582 clips, 85 genres) and 15 of 20 in
levo2-vocals-gender-rebalance, but those corpora use short tag-style descriptions
(“male, rock, sixties surf pop lead vocal, …”) — 210 and 60 distinct
prompts respectively, and zero of them match any of the 20 prompts used here. Under a
rule of identical prompts to both models, genre-matched is not prompt-matched, so none of them
were eligible.
Generation
MiniMax Music 3 — the diffusers modular pipeline. The model card's three
runtimes are SGLang-Omni, diffusers and ComfyUI; the sglang-omni build present on this system
has no minimax_music3 model, so diffusers was the route taken. It lives in an
unmerged PR, installed with --no-deps into a directory on
PYTHONPATH so no existing environment was modified.
One GH200 per shard, four shards, bf16, peak 22.8 GiB — it fits the 24 GB card the model card claims. 61 s of wall clock per clip (20 GPU-minutes for all 20), 44.1 kHz stereo, audio_duration=45, 30 flow-matching steps per chunk. The language model emits an end-of-audio token when it has finished the lyric, so the renders run 15–30 s (mean 22.4 s) rather than the full 45.
LeVo 2 / SongGeneration-v2-large — gen_type='vocal' for the main
arm and gen_type='mixed' for the control, otherwise identical settings
(temperature 0.9, top-k 50, cfg 2.0, 30 s). 50 s per clip, 48 kHz stereo, renders 20–30 s (mean 28.2 s). The two arms therefore differ in mean clip length by about 6 s, which is stated because a shorter clip gives an instrument less time to appear; the objective accompaniment fraction is an energy ratio and is not affected by it.
Both arms: one seed per prompt, the same seed (4242) for every clip, 30 s kept.
What it took to run MiniMax Music 3 here
Recorded because “it runs” was not a given: aarch64 GH200, torch already installed,
compute nodes with no internet. Nothing was pip installed without
--no-deps, and no existing environment was modified.
Weights — 57 GB in the repo; only the 28 GB the diffusers path needs was
fetched (the 9.8 GB flowmatching_vae.pth and the 48-shard qwen_7B/
tree serve the SGLang runtime). Fetched on the login node, 32 files in 29 s.
diffusers — MiniMax Music 3 is in an unmerged PR
(huggingface/diffusers@dafe3733). Installed with
uv pip install --no-deps --target song/dfmm and put on PYTHONPATH, so
env_song's own diffusers 0.39.0 is untouched. That is the only thing installed for
this work.
The hub id does not work offline.modular_model_index.json names the
repo id for every component, so with a warm cache and HF_HUB_OFFLINE=1 the loader
still calls the Hub API and fails. Fixed by materialising a directory of symlinks with the
component paths rewritten to point at it (mm_localise.py).
The tokenizer folder is incomplete for the class the pipeline declares. The pipeline
asks for the slow Qwen2Tokenizer, which wants vocab.json +
merges.txt; tokenizer/ ships only tokenizer.json, and the
protobuf fallback is not installed. The repo's other tokenizer copy has both files, and it was
checked to be the same tokenizer — identical vocab, merges and added tokens, differing
only in serialised bytes — before being linked in.
One import from the future.diffusers/pipelines/pipeline_utils.py
imports get_cached_repo_tree / CachedRepoTreeNotFoundError, which
postdate huggingface_hub 0.36.2. Shimmed in-process to raise the "nothing cached" error rather
than upgrading huggingface_hub underneath ACE-Step.
It then just works. bf16, peak 22.8 GiB, no flash-attention wheel needed, no
custom kernels, 60 s per clip on one GH200. The vocoder reports 44.1 kHz, not the
32 kHz the model card states.
Verifying the LeVo 2 baseline rather than assuming it
The prior figures for LeVo 2 on this task were singing quality 4.40, one-voice compliance
4.07, no change across stem separation, and a measured accompaniment fraction of 0.0037.
Two of those reproduce and two do not, for reasons worth stating:
No change across stem separation: reproduced exactly. LeVo's one-voice score is
4.28 on the mix and 4.28 on the stem — a delta of
+0.00. There is nothing for a separator to take away.
Accompaniment fraction: lower still, 0.0000 (max over 20 clips
0.0002). Measured here as the separator's own energy split rather than by the
earlier method, so the two numbers are not the same statistic; both say the same thing.
Singing quality 3.77, not 4.40, and one-voice 4.28, not 4.07.
Neither is a contradiction: this is a different rubric with different anchors, a different
prompt, one seed per genre, and a genre set chosen for spread rather than for what LeVo does
well. The prior numbers are not a baseline this run can be scored against, which is why they
are not used as one anywhere above.
Judging
Gemini 3.6 Flash over the native generateContent endpoint with
thinkingConfig: {"thinkingLevel": "low"} —
{"thinkingBudget": 0} returns HTTP 400 for this model on this endpoint,
verified against the live API before the run. Audio at 128 kbps, because one of the axes
is audio quality and the codec must not be the thing being graded.
Blind by construction. Each request carries exactly one clip plus the prompt and
lyrics it was generated from. No model name, no file name, no arm label, and never two clips in
one request, so there is nothing in the request that could identify the generator; the queue is
shuffled as well, so dispatch order carries no signal either. Each clip was judged twice, in two
independent passes with separate caches.
All 20, side by side
Each row shows the prompt exactly as both models received it, both takes, the five per-clip scores (mean of the two judging passes), the accompaniment fraction, and the judge's own one-line description of what it heard. The note under each row is generated from the numbers, not written by hand.
1. Sámi joikjoik · world · adult · any · stands for 8 genres
a cappella solo voice, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; sami joik, cyclical wordless vocable chant, guttural and hypnotic; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 2.0I 2.0G 1.0O 1.0
vocal stem: S 3.0 V 5.0 I 4.5 G 1.0 O 1.5
accompaniment fraction 0.6143
“A male lead voice singing folk-pop accompanied by acoustic guitar, subtle ambient synth pads, percussion snaps, and layering vocal harmonies.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 5.0G 1.0O 2.0
vocal stem: S 4.0 V 5.0 I 5.0 G 1.0 O 2.0
accompaniment fraction 0.0000
“a single female voice singing a soft folk-pop melody with no instruments or backing voices”
LeVo 2 (gen_type=mixed)
S 4.0V 5.0I 2.0G 1.0O 1.0
vocal stem: S 4.0 V 5.0 I 5.0 G 1.0 O 2.0
accompaniment fraction 0.7034
“A female voice singing a sweet folk/pop melody accompanied by a prominent acoustic ukulele or guitar and light percussion.”
Overall 1.0 vs 2.0: LeVo ahead by 1.0, wider than the 0.45-point noise floor. The widest axis is one voice alone (2.0 vs 5.0). Accompaniment fraction 0.614 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 3.0, so what the judge heard was instrumental bleed.
2. Teen poppop_teen · pop · teen · female · stands for 6 genres
a cappella solo voice, female, a teenager; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; teen pop, bright youthful thin-bodied voice, eager and unweathered; dry close-mic, no reverb
MiniMax Music 3
S 2.5V 4.0I 3.5G 4.0O 2.0
vocal stem: S 3.0 V 4.0 I 2.0 G 3.0 O 2.0
accompaniment fraction 0.0990
“A single young female voice singing dry and completely unaccompanied, though with heavy digital glitchey vocal artifacts.”
LeVo 2 (gen_type=vocal)
S 4.0V 4.0I 5.0G 5.0O 4.5
vocal stem: S 4.5 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.0000
“a single teen female voice singing alone in the verses, joined by layered harmonies/doubling in the chorus with no instruments present”
LeVo 2 (gen_type=mixed)
S 4.0V 3.5I 2.0G 4.0O 2.0
vocal stem: S 4.0 V 2.0 I 5.0 G 4.5 O 3.0
accompaniment fraction 0.7785
“A female teenage voice singing with a full instrumental arrangement including guitar, synths, and percussion.”
Overall 2.0 vs 4.5: LeVo ahead by 2.5, wider than the 0.45-point noise floor. The widest axis is singing quality (2.5 vs 4.0). Accompaniment fraction 0.099 vs 0.000. Stripping the instruments moves MiniMax's one-voice score by +0.0 — barely, so the extra voice is a real second singer, not an instrument.
3. Punkpunk · rock · young_adult · male · stands for 4 genres
a cappella solo voice, male, in their twenties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; punk, shouted snarling barked delivery, aggressive and untrained; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 5.0I 2.0G 1.0O 1.0
vocal stem: S 3.0 V 5.0 I 4.0 G 1.0 O 1.5
accompaniment fraction 0.4362
“A single male voice singing a soft folk-pop melody accompanied by strummed acoustic guitar and ambient synth pads.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 4.5G 4.5O 4.5
vocal stem: S 4.0 V 5.0 I 5.0 G 4.5 O 4.5
accompaniment fraction 0.0002
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 3.0V 5.0I 2.0G 3.0O 1.5
vocal stem: S 4.5 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.7231
“A male voice singing with an acoustic guitar and backing drums/bass instrumental track.”
Overall 1.0 vs 4.5: LeVo ahead by 3.5, wider than the 0.45-point noise floor. The widest axis is genre match (1.0 vs 4.5). Accompaniment fraction 0.436 vs 0.000.
4. Indie popindie_pop · pop · young_adult · any · stands for 1 genre
a cappella solo voice, in their twenties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; indie pop, slightly nasal untrained charm, close-mic breathiness, casual; dry close-mic, no reverb
MiniMax Music 3
S 4.0V 5.0I 3.5G 4.0O 3.5
vocal stem: S 3.5 V 5.0 I 5.0 G 4.5 O 4.0
accompaniment fraction 0.1964
“A young female voice singing accompanied by an acoustic guitar.”
LeVo 2 (gen_type=vocal)
S 4.5V 5.0I 5.0G 5.0O 5.0
vocal stem: S 5.0 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.0000
“A single young female voice singing completely unaccompanied with no instruments or extra vocal layers present.”
LeVo 2 (gen_type=mixed)
S 4.0V 2.0I 2.0G 3.5O 1.0
vocal stem: S 3.5 V 2.0 I 5.0 G 5.0 O 3.0
accompaniment fraction 0.7879
“A female indie pop voice singing over a full instrumental track of drums, bass, and guitar, with layered harmonized vocals joining in the chorus.”
Overall 3.5 vs 5.0: LeVo ahead by 1.5, wider than the 0.45-point noise floor. The widest axis is no instruments (3.5 vs 5.0). Accompaniment fraction 0.196 vs 0.000.
5. Afrobeatsafrobeats · urban · young_adult · male · stands for 1 genre
a cappella solo voice, male, in their twenties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; afrobeats, lilting melodic pidgin-inflected phrasing, relaxed buoyant swing; dry close-mic, no reverb
MiniMax Music 3
S 2.5V 3.5I 3.0G 2.0O 2.0
vocal stem: S 1.0 V 5.0 I 5.0 G 1.0 O 1.0
accompaniment fraction 0.5643
“a male lead voice accompanied by finger snaps, subtle pad/bass tones, and layered vocal harmonies.”
LeVo 2 (gen_type=vocal)
S 4.0V 4.0I 5.0G 4.0O 3.5
vocal stem: S 4.0 V 5.0 I 4.0 G 4.0 O 3.5
accompaniment fraction 0.0000
“a single young male voice singing with vocal pops, snaps, rhythmic vocal noises, and faint background vocal harmonizations/doubling”
LeVo 2 (gen_type=mixed)
S 4.0V 2.5I 2.0G 4.5O 2.0
vocal stem: S 4.5 V 4.5 I 5.0 G 4.5 O 4.5
accompaniment fraction 0.5730
“A male voice singing Afrobeats with backing vocal ad-libs, acoustic guitar, bass, shaker, and percussion.”
Overall 2.0 vs 3.5: LeVo ahead by 1.5, wider than the 0.45-point noise floor. The widest axis is no instruments (3.0 vs 5.0). Accompaniment fraction 0.564 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 1.5, so what the judge heard was instrumental bleed.
6. Gothic metal sopranometal_gothic_female · metal · adult · female · stands for 2 genres
a cappella solo voice, female, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; gothic symphonic metal soprano, operatic vibrato over heaviness, ethereal; dry close-mic, no reverb
MiniMax Music 3
S 3.5V 3.5I 2.0G 3.5O 2.0
vocal stem: S 4.0 V 5.0 I 5.0 G 3.0 O 4.0
accompaniment fraction 0.7159
“A single operatic female soprano singing alone initially, joined by heavy metal percussion and drums halfway through.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 5.0G 2.0O 3.0
vocal stem: S 3.0 V 5.0 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 3.5V 2.0I 2.0G 1.0O 1.0
vocal stem: S 3.0 V 1.5 I 5.0 G 1.5 O 1.5
accompaniment fraction 0.4651
“A female voice singing acoustic folk-pop accompanied by acoustic guitar strumming and backing harmonies.”
Overall 2.0 vs 3.0: LeVo ahead by 1.0, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.716 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 1.5, so what the judge heard was instrumental bleed.
7. Chorister solochoirboy_solo · classical · child · male · stands for 1 genre
a cappella solo voice, male, a child of about eight; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; solo chorister boy, pure vibrato-free treble, echoing cathedral clarity; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 5.0I 2.0G 1.5O 1.0
vocal stem: S 2.5 V 5.0 I 3.0 G 2.0 O 2.0
accompaniment fraction 0.4206
“An adult male singer performing with acoustic guitar accompaniment.”
LeVo 2 (gen_type=vocal)
S 3.5V 5.0I 5.0G 4.0O 4.0
vocal stem: S 3.5 V 5.0 I 5.0 G 4.0 O 4.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 4.0V 4.0I 2.0G 0.5O 1.0
vocal stem: S 3.0 V 5.0 I 5.0 G 4.0 O 3.5
accompaniment fraction 0.5303
“I hear an adult female singer accompanied by a full pop/soul instrumental arrangement with drums, bass, and electric guitar.”
Overall 1.0 vs 4.0: LeVo ahead by 3.0, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.421 vs 0.000.
8. Power metalmetal_power · metal · adult · male · stands for 1 genre
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; power metal, soaring operatic high tenor, sustained ringing screams; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 5.0I 2.0G 1.0O 1.0
vocal stem: S 2.0 V 2.5 I 5.0 G 1.5 O 1.5
accompaniment fraction 0.2030
“A male voice singing with prominent acoustic guitar accompaniment, fingerpicking, and rhythmic tapping sounds.”
LeVo 2 (gen_type=vocal)
S 3.0V 2.5I 2.0G 2.0O 1.5
vocal stem: S 3.0 V 2.0 I 5.0 G 2.5 O 2.0
accompaniment fraction 0.0000
“A male vocal singing in a folk/rock style, accompanied by acoustic guitar strumming and brief backing harmonies in the chorus.”
LeVo 2 (gen_type=mixed)
S 4.0V 2.0I 2.0G 2.0O 1.0
vocal stem: S 3.0 V 2.5 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.5648
“A male voice singing with prominent acoustic guitar strums and double-tracked/harmony voices joining in the chorus.”
Overall 1.0 vs 1.5: LeVo ahead by 0.5, wider than the 0.45-point noise floor. The widest axis is one voice alone (5.0 vs 2.5). Accompaniment fraction 0.203 vs 0.000.
9. Boy-band tenorpop_boyband · pop · teen · male · stands for 1 genre
a cappella solo voice, male, a teenager; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; boy-band lead, light high tenor, earnest and smooth, adolescent brightness; dry close-mic, no reverb
MiniMax Music 3
S 2.5V 2.0I 2.0G 3.0O 1.0
vocal stem: S 1.0 V 5.0 I 5.0 G 1.5 O 1.0
accompaniment fraction 0.3268
“A male vocalist accompanied by an acoustic guitar, percussion snaps, and stacked backing vocals.”
LeVo 2 (gen_type=vocal)
S 4.5V 5.0I 5.0G 5.0O 5.0
vocal stem: S 4.5 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 4.5V 5.0I 2.0G 4.5O 2.0
vocal stem: S 5.0 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.4698
“A single young male voice singing lead accompanied by an audible electric piano track.”
Overall 1.0 vs 5.0: LeVo ahead by 4.0, wider than the 0.45-point noise floor. The widest axis is one voice alone (2.0 vs 5.0). Accompaniment fraction 0.327 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 3.0, so what the judge heard was instrumental bleed.
10. Patter songpatter_song · theatre · adult · male · stands for 2 genres
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; comic patter song, extremely fast precise consonants, breathless wit; dry close-mic, no reverb
MiniMax Music 3
S 2.5V 1.0I 2.0G 0.5O 1.0
vocal stem: S 2.0 V 0.5 I 5.0 G 0.5 O 1.0
accompaniment fraction 0.4908
“A male lead voice accompanied by harmonizing backing vocals and a sustained synth pad.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 4.0G 2.0O 2.5
vocal stem: S 4.0 V 5.0 I 5.0 G 2.5 O 3.0
accompaniment fraction 0.0000
“a single male voice singing accompanied by rhythmic finger snaps”
LeVo 2 (gen_type=mixed)
S 3.5V 5.0I 2.0G 1.5O 1.0
vocal stem: S 3.5 V 5.0 I 4.5 G 3.0 O 3.0
accompaniment fraction 0.6163
“A male voice singing a mid-tempo pop melody accompanied by an acoustic guitar, bass, and percussion beat.”
Overall 1.0 vs 2.5: LeVo ahead by 1.5, wider than the 0.45-point noise floor. The widest axis is one voice alone (1.0 vs 5.0). Accompaniment fraction 0.491 vs 0.000. Stripping the instruments moves MiniMax's one-voice score by -0.5 — barely, so the extra voice is a real second singer, not an instrument.
11. Bel cantobel_canto · classical · adult · any · stands for 1 genre
a cappella solo voice, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; bel canto, effortless even legato across registers, elegant messa di voce swells; dry close-mic, no reverb
MiniMax Music 3
S 1.5V 3.0I 2.0G 1.0O 1.0
vocal stem: S 1.0 V 3.5 I 5.0 G 1.0 O 1.0
accompaniment fraction 0.4579
“A glitchy, unstable voice singing fragmented lines accompanied by acoustic guitar chords and subtle bass towards the end.”
LeVo 2 (gen_type=vocal)
S 3.5V 5.0I 3.5G 1.5O 1.5
vocal stem: S 3.0 V 5.0 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.0000
“a single female voice singing unaccompanied with rhythmic finger snaps or quiet clicking in the background”
LeVo 2 (gen_type=mixed)
S 3.5V 3.5I 2.0G 1.0O 1.0
vocal stem: S 3.5 V 5.0 I 5.0 G 3.5 O 3.5
accompaniment fraction 0.5160
“A female pop vocalist singing accompanied by an upbeat synth, bass, and rhythmic electronic drums arrangement.”
Overall 1.0 vs 1.5: LeVo ahead by 0.5, wider than the 0.45-point noise floor. The widest axis is one voice alone (3.0 vs 5.0). Accompaniment fraction 0.458 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 0.5, so what the judge heard was instrumental bleed.
12. Cabaretcabaret · theatre · adult · any · stands for 2 genres
a cappella solo voice, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; cabaret, arch knowing delivery, spoken asides sliding into sung line; dry close-mic, no reverb
MiniMax Music 3
S 3.5V 3.5I 2.0G 3.0O 1.0
vocal stem: S 1.5 V 5.0 I 5.0 G 2.5 O 1.5
accompaniment fraction 0.4872
“A male voice singing in a theatrical style accompanied by an acoustic guitar throughout.”
LeVo 2 (gen_type=vocal)
S 3.0V 5.0I 5.0G 4.5O 3.5
vocal stem: S 3.5 V 2.0 I 5.0 G 4.5 O 3.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 4.0V 2.5I 2.5G 4.5O 2.0
vocal stem: S 4.5 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.6694
“A female cabaret vocal with percussive finger snaps, vocal bass percussion, and layered backing vocal harmonies in the chorus.”
Overall 1.0 vs 3.5: LeVo ahead by 2.5, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.487 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 1.5, so what the judge heard was instrumental bleed.
13. Forróforro · world · adult · male · stands for 1 genre
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; forro, warm northeastern brazilian accent, rustic bouncing delivery; dry close-mic, no reverb
MiniMax Music 3
S 2.0V 4.5I 4.0G 2.0O 2.0
vocal stem: S 2.0 V 4.0 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.3016
“A high, heavily artifacted and glitching male vocal with audible pitch-shifting artifacts and no instruments.”
LeVo 2 (gen_type=vocal)
S 4.0V 3.0I 5.0G 2.5O 2.5
vocal stem: S 3.5 V 3.0 I 5.0 G 3.5 O 3.0
accompaniment fraction 0.0000
“A male voice singing alone, joined by self-harmony and layering in the chorus, with no musical instruments present.”
LeVo 2 (gen_type=mixed)
S 4.0V 2.0I 1.5G 3.5O 1.5
vocal stem: S 4.0 V 3.0 I 5.0 G 2.5 O 2.5
accompaniment fraction 0.3189
“A male lead voice singing with a upbeat forró style acoustic bass, accordion, and percussive instrumentation in the background, along with backing vocal harmonies during the chorus.”
Overall 2.0 vs 2.5: LeVo ahead by 0.5, wider than the 0.45-point noise floor. The widest axis is singing quality (2.0 vs 4.0). Accompaniment fraction 0.302 vs 0.000.
14. Saetaflamenco_sae · world · adult · female · stands for 3 genres
a cappella solo voice, female, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; spanish saeta, unaccompanied processional cry from a balcony, raw and piercing; dry close-mic, no reverb
MiniMax Music 3
S 4.0V 5.0I 4.0G 5.0O 4.5
vocal stem: S 4.0 V 5.0 I 5.0 G 4.0 O 4.0
accompaniment fraction 0.1767
“nothing but a single voice”
LeVo 2 (gen_type=vocal)
S 3.5V 5.0I 5.0G 1.0O 2.0
vocal stem: S 3.5 V 5.0 I 5.0 G 1.0 O 2.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 3.0V 5.0I 2.0G 0.5O 1.0
vocal stem: S 4.0 V 5.0 I 5.0 G 2.0 O 2.5
accompaniment fraction 0.6737
“A female singing accompanied by rhythmic claps, bass guitar, and a ukuele/acoustic guitar arrangement.”
Overall 4.5 vs 2.0: MiniMax ahead by 2.5, wider than the 0.45-point noise floor. The widest axis is genre match (5.0 vs 1.0). Accompaniment fraction 0.177 vs 0.000.
15. Rancheraranchera · world · adult · male · stands for 2 genres
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; mexican ranchera, huge open-throated grito, proud sustained belt; dry close-mic, no reverb
MiniMax Music 3
S 2.0V 3.5I 2.0G 1.0O 1.0
vocal stem: S 2.0 V 3.0 I 4.5 G 1.0 O 1.0
accompaniment fraction 0.6745
“A single male voice singing softly, accompanied by acoustic guitar strumming, bass, and atmospheric synth pads.”
LeVo 2 (gen_type=vocal)
S 3.0V 2.0I 5.0G 2.0O 2.0
vocal stem: S 2.5 V 3.5 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.0000
“A male voice starts alone singing in a gentle pop/folk style, but is then joined by stacked harmonizing voices that create a choir effect in the chorus, with no musical instruments present.”
LeVo 2 (gen_type=mixed)
S 3.5V 5.0I 2.0G 1.5O 1.0
vocal stem: S 2.5 V 2.0 I 5.0 G 1.5 O 1.5
accompaniment fraction 0.6325
“A male voice singing with an acoustic bass and acoustic guitar backing track.”
Overall 1.0 vs 2.0: LeVo ahead by 1.0, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.675 vs 0.000. Stripping the instruments moves MiniMax's one-voice score by -0.5 — barely, so the extra voice is a real second singer, not an instrument.
16. Quiet stormquiet_storm · soul · adult · male · stands for 1 genre
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; quiet storm R&B, smooth low seductive croon, silky sustained falsetto; dry close-mic, no reverb
MiniMax Music 3
S 4.0V 3.0I 2.0G 4.0O 2.0
vocal stem: S 2.0 V 5.0 I 2.5 G 2.0 O 1.5
accompaniment fraction 0.4469
“A male R&B voice singing accompanied by finger snaps, synth pads, guitar chords, and light background harmony vocals.”
LeVo 2 (gen_type=vocal)
S 4.0V 3.5I 5.0G 5.0O 3.5
vocal stem: S 4.0 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.0000
“A single smooth male voice singing a cappella with a brief self-harmony/doubling layer appearing on the final phrase, and no instruments present.”
LeVo 2 (gen_type=mixed)
S 4.0V 3.0I 2.0G 3.5O 1.5
vocal stem: S 3.5 V 5.0 I 5.0 G 4.5 O 4.0
accompaniment fraction 0.5965
“A smooth male voice singing R&B accompanied by a full instrumental track including drums, bass, and guitar.”
Overall 2.0 vs 3.5: LeVo ahead by 1.5, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.447 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 2.0, so what the judge heard was instrumental bleed.
17. Delta bluesblues_delta · blues · older · male · stands for 2 genres
a cappella solo voice, male, over sixty; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; delta blues, rough weathered voice, bent notes, moaning field-holler inflection; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 3.5I 2.0G 1.5O 1.0
vocal stem: S 3.0 V 4.0 I 5.0 G 3.5 O 3.0
accompaniment fraction 0.3581
“A young male voice singing accompanied by acoustic guitar strumming and foot stomps/claps, with layered backing vocals joining in on the chorus.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 5.0G 2.5O 3.0
vocal stem: S 3.5 V 4.0 I 5.0 G 2.0 O 2.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 4.0V 5.0I 2.0G 1.0O 1.0
vocal stem: S 2.5 V 5.0 I 5.0 G 2.5 O 2.5
accompaniment fraction 0.3740
“A female singing voice with acoustic bass, finger snapping, and light drums accompanying the melody.”
Overall 1.0 vs 3.0: LeVo ahead by 2.0, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.358 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 0.5, so what the judge heard was instrumental bleed.
18. Dancehalldancehall · urban · adult · male · stands for 2 genres
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; dancehall, rapid patois toasting sliding into melody, rhythmic and percussive; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 5.0I 2.0G 1.0O 1.0
vocal stem: S 2.0 V 2.0 I 5.0 G 1.0 O 1.0
accompaniment fraction 0.3849
“A male voice singing accompanied by an acoustic guitar.”
LeVo 2 (gen_type=vocal)
S 3.0V 1.5I 4.0G 0.0O 0.5
vocal stem: S 3.5 V 3.5 I 5.0 G 0.5 O 1.0
accompaniment fraction 0.0000
“A female vocalist singing a folk ballad with background synth pads and multi-part harmonic backing vocals.”
LeVo 2 (gen_type=mixed)
S 3.0V 2.5I 2.0G 1.5O 1.0
vocal stem: S 2.5 V 2.5 I 5.0 G 1.0 O 1.0
accompaniment fraction 0.7152
“A female synth-pop singing voice accompanied by a prominent synth bass, electronic percussion beat, and vocal harmonies.”
Overall 1.0 vs 0.5: MiniMax ahead by 0.5, wider than the 0.45-point noise floor. The widest axis is one voice alone (5.0 vs 1.5). Accompaniment fraction 0.385 vs 0.000.
19. Pure falsettofalsetto_pure · functional · adult · male · stands for 1 genre
a cappella solo voice, male, in their thirties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; pure falsetto, light hooty unsupported head voice, fragile; dry close-mic, no reverb
MiniMax Music 3
S 3.0V 2.5I 2.0G 2.0O 1.0
vocal stem: S 1.5 V 4.5 I 3.5 G 2.5 O 1.5
accompaniment fraction 0.0900
“I hear a male voice accompanied by an acoustic guitar and layered vocal harmonies.”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 5.0G 5.0O 5.0
vocal stem: S 4.0 V 2.5 I 5.0 G 4.0 O 3.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 4.0V 2.5I 2.0G 1.5O 1.0
vocal stem: S 3.0 V 2.0 I 5.0 G 3.5 O 2.0
accompaniment fraction 0.3843
“A male falsetto singing with acoustic guitar, light percussion, and backing vocal harmonies.”
Overall 1.0 vs 5.0: LeVo ahead by 4.0, wider than the 0.45-point noise floor. The widest axis is no instruments (2.0 vs 5.0). Accompaniment fraction 0.090 vs 0.000. Stripping the instruments raises MiniMax's one-voice score by 2.0, so what the judge heard was instrumental bleed.
20. Britpopbritpop · rock · young_adult · male · stands for 1 genre
a cappella solo voice, male, in their twenties; unaccompanied, no instruments, no backing track, no second voice, no harmony, no choir; britpop, adenoidal regional-accented drawl, cocky flat delivery; dry close-mic, no reverb
MiniMax Music 3
S 2.5V 5.0I 5.0G 3.5O 2.5
vocal stem: S 2.5 V 5.0 I 5.0 G 3.5 O 2.5
accompaniment fraction 0.0578
“nothing but a single male voice”
LeVo 2 (gen_type=vocal)
S 4.0V 5.0I 5.0G 5.0O 5.0
vocal stem: S 5.0 V 5.0 I 5.0 G 5.0 O 5.0
accompaniment fraction 0.0000
“nothing but a single voice”
LeVo 2 (gen_type=mixed)
S 3.0V 2.5I 2.0G 3.5O 1.5
vocal stem: S 3.0 V 2.0 I 5.0 G 3.0 O 2.0
accompaniment fraction 0.5304
“A male lead voice singing with acoustic guitar, light percussion, and layered backing vocals during the chorus.”
Overall 2.5 vs 5.0: LeVo ahead by 2.5, wider than the 0.45-point noise floor. The widest axis is singing quality (2.5 vs 4.0). Accompaniment fraction 0.058 vs 0.000.
Built 2026-08-14. 60 clips generated, 60 judged twice on the mix and 60 on the vocal stem. Scripts: song/scripts/mm_prompts.py, gen_minimax.py, gen_levo.py, mm_stems.py, mm_judge.py, mm_agg.py, mm_page.py.