Every step, adapter, prompt and setting actually used in the two-profile demo run, with the measured result of each.
k325_age3_bg1 —
Velvet Sage Baritone,
Male,
Late 40s to 60s
variant chosencbx (best DNSMOS of orig/sidon/cbx)
groups808
candidates / group200
profilesA = documented best practice · B = identical +
laion/moss-mediathek-emotion-lora :: r64_e2 merged at
25 %
total generations323,200
hardware16 nodes × 4 GH200, one worker per GPU, 32 shards × 2
profiles, batch 32
base modellaion/moss-tts-local-transformer-4.55b-voice-acting-v2, bf16, sdpa
Profile B differs from Profile A in exactly one respect: the Mediathek adapter
is added to every set_active spec at 0.25. Same seeds, same
captions, same sampling, same text, same reference.
| block | groups | candidates | mean gen s/group | mean score s/group | gen / GPU-hour | empty decodes |
|---|---|---|---|---|---|---|
| edge | 56 | 11200 | 329 | 13 | 2,106 | 0.00% |
| emotion | 640 | 128000 | 331 | 13 | 2,092 | 0.01% |
| explicit | 4 | 800 | 312 | 12 | 2,220 | 0.00% |
| sports | 4 | 800 | 157 | 12 | 4,263 | 0.00% |
| voicenet | 912 | 182400 | 322 | 13 | 2,146 | 0.17% |
| all | 1616 | 323200 | 2,126 | 0.10% |
| check | result |
|---|---|
| EN/DE cells that share a text slot | 404 / 404 — 0 mismatched |
| distinct text slots per emotion (must be exactly 1 across its 8 groups) | [1] |
| groups in the matrix | 808 |
| candidates stored per group | 200 — every one of them, ranked |
A worked example: the four conditions of Fear and both languages
all draw the same slot emo::Fear, so the EN and DE takes differ only in
language.
| EN | The infusion has a mild taste and goes well with herbs, tea, or mate tea. |
| DE | Der Aufguss hat einen milden Geschmack und passt gut zu Kräutern, Tee oder Mate Tee. |
| block | groups | score | spk_sim | WER | genuineness | blend | strength | naturalness | dur s | DFLU | words/s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| edge | 28 | 0.523 | 0.459 | 0.155 | 1.85 | 7.77 | 2.92 | 0.732 | 8.0 | 2.42 | 1.89 |
| emotion | 320 | 0.592 | 0.490 | 0.074 | 0.74 | 3.22 | 0.60 | 0.592 | 9.0 | 1.07 | 2.70 |
| explicit | 2 | 0.571 | 0.553 | 0.215 | 1.50 | 2.52 | 2.27 | 0.731 | 9.7 | 2.34 | 1.45 |
| sports | 2 | 0.578 | 0.439 | 0.094 | 3.99 | 4.76 | 2.12 | 0.600 | 4.9 | 2.86 | 6.39 |
| voicenet | 456 | 0.643 | 0.535 | 0.099 | 1.12 | 2.98 | 0.78 | 0.667 | 10.1 | 1.58 | 2.36 |
| block | groups | score | spk_sim | WER | genuineness | blend | strength | naturalness | dur s | DFLU | words/s |
|---|---|---|---|---|---|---|---|---|---|---|---|
| edge | 28 | 0.515 | 0.460 | 0.153 | 1.89 | 7.40 | 2.80 | 0.766 | 7.9 | 2.36 | 1.90 |
| emotion | 320 | 0.591 | 0.501 | 0.076 | 0.78 | 3.20 | 0.57 | 0.591 | 9.0 | 1.11 | 2.71 |
| explicit | 2 | 0.555 | 0.511 | 0.249 | 1.55 | 2.82 | 1.96 | 0.775 | 8.4 | 2.17 | 1.61 |
| sports | 2 | 0.568 | 0.497 | 0.016 | 3.21 | 4.97 | 2.14 | 0.609 | 5.7 | 2.60 | 5.76 |
| voicenet | 456 | 0.643 | 0.542 | 0.101 | 1.14 | 2.89 | 0.86 | 0.662 | 10.7 | 1.63 | 2.30 |
The comparison the design exists to make. Read alongside the published pooled figures for the same four conditions on six other voices (A emo 0.437 / genu 0.362 / spk 0.569 · B 0.273 / 0.400 / 0.617 · C 0.294 / 0.265 / 0.559 · D 0.185 / 0.326 / 0.617), where the headline was that intensity, not containment, breaks the clone.
| condition | groups | emotion strength | genuineness | blend | spk_sim | VULN | WER | naturalness | below floor |
|---|---|---|---|---|---|---|---|---|---|
| A — Intense, free | 80 | 0.771 | 0.955 | 3.534 | 0.422 | 1.909 | 0.078 | 0.648 | 19 / 80 |
| B — Moderate, free | 80 | 0.480 | 0.622 | 3.223 | 0.537 | 1.972 | 0.073 | 0.580 | 0 / 80 |
| C — Intense, contained | 80 | 0.761 | 0.731 | 3.195 | 0.471 | 1.563 | 0.074 | 0.575 | 12 / 80 |
| D — Moderate, contained | 80 | 0.445 | 0.637 | 2.920 | 0.533 | 1.661 | 0.070 | 0.563 | 0 / 80 |
Containment contrast (C,D − A,B): VULN -0.329, emotion strength -0.022, genuineness -0.104.
Intensity contrast (A,C − B,D): emotion strength +0.304, speaker similarity -0.088, below-floor rate 19.4 % vs 0.0 %.
The EN/DE pairing is the point of the corpus, so the two halves are reported separately. They use the same sentence, the same voice, the same adapters and the same sampling.
| language | groups | WER | spk_sim | genuineness | blend | emotion strength | dur s | words/s | naturalness |
|---|---|---|---|---|---|---|---|---|---|
| EN | 404 | 0.083 | 0.607 | 0.983 | 3.367 | 1.434 | 8.3 | 2.79 | 0.689 |
| DE | 404 | 0.100 | 0.422 | 1.015 | 3.118 | 1.476 | 10.8 | 2.18 | 0.589 |
Paired on the rank-0 take of every group that both profiles produced. t is
a paired t-statistic on the per-group difference; with hundreds of groups, |t| > 3
is a real effect and |t| < 2 is not.
| block | n | score | genuineness | blend | strength_raw | wer | spk_sim | naturalness | quality | dur | dflu | wps | vuln |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ALL | 808 | 0.619→0.617 -0.001 · B wins 49 % · t=-0.4 | 0.999→1.030 +0.030 · B wins 48 % · t=+1.5 | 3.243→3.176 -0.067 · B wins 49 % · t=-1.0 | 0.791→0.817 +0.027 · B wins 45 % · t=+1.1 | 0.091→0.093 +0.001 · B wins 21 % · t=+0.7 | 0.515→0.523 +0.008 · B wins 54 % · t=+2.6 | 0.639→0.638 -0.001 · B wins 47 % · t=-0.2 | 3.400→3.401 +0.000 · B wins 53 % · t=+0.0 | 9.554→9.885 +0.330 · B wins 50 % · t=+2.6 | 1.414→1.451 +0.037 · B wins 52 % · t=+1.4 | 2.488→2.454 -0.034 · B wins 46 % · t=-2.1 | 1.986→1.975 -0.011 · B wins 50 % · t=-0.4 |
| edge | 28 | 0.523→0.515 -0.008 · B wins 36 % · t=-0.8 | 1.853→1.885 +0.032 · B wins 50 % · t=+0.3 | 7.771→7.397 -0.373 · B wins 36 % · t=-1.4 | 2.917→2.797 -0.121 · B wins 32 % · t=-1.1 | 0.155→0.153 -0.002 · B wins 14 % · t=-0.2 | 0.459→0.460 +0.001 · B wins 50 % · t=+0.1 | 0.732→0.766 +0.034 · B wins 50 % · t=+0.8 | 3.340→3.349 +0.008 · B wins 50 % · t=+0.5 | 8.000→7.906 -0.094 · B wins 36 % · t=-0.3 | 2.422→2.364 -0.058 · B wins 50 % · t=-0.2 | 1.892→1.900 +0.008 · B wins 64 % · t=+0.1 | 3.852→3.923 +0.071 · B wins 43 % · t=+0.3 |
| emotion | 320 | 0.592→0.591 -0.001 · B wins 48 % · t=-0.2 | 0.736→0.782 +0.046 · B wins 49 % · t=+1.5 | 3.218→3.199 -0.019 · B wins 48 % · t=-0.2 | 0.596→0.570 -0.026 · B wins 32 % · t=-0.9 | 0.074→0.076 +0.002 · B wins 21 % · t=+0.7 | 0.490→0.501 +0.011 · B wins 55 % · t=+2.2 | 0.592→0.591 -0.001 · B wins 50 % · t=-0.1 | 3.424→3.421 -0.003 · B wins 53 % · t=-0.7 | 8.965→8.966 +0.002 · B wins 49 % · t=+0.0 | 1.074→1.108 +0.034 · B wins 53 % · t=+1.0 | 2.702→2.710 +0.007 · B wins 45 % · t=+0.3 | 1.776→1.782 +0.005 · B wins 50 % · t=+0.2 |
| explicit | 2 | 0.571→0.555 -0.016 · B wins 0 % · t=-2.3 | 1.504→1.546 +0.042 · B wins 50 % · t=+0.1 | 2.521→2.821 +0.300 · B wins 50 % · t=+0.2 | 2.267→1.964 -0.303 · B wins 0 % · t=-1.0 | 0.215→0.249 +0.033 · B wins 50 % · t=+1.0 | 0.553→0.511 -0.042 · B wins 0 % · t=-1.2 | 0.731→0.775 +0.043 · B wins 50 % · t=+1.0 | 3.418→3.459 +0.041 · B wins 100 % · t=+3.0 | 9.720→8.400 -1.320 · B wins 0 % · t=-4.7 | 2.345→2.173 -0.171 · B wins 50 % · t=-0.5 | 1.448→1.614 +0.166 · B wins 100 % · t=+3.3 | 2.374→2.806 +0.432 · B wins 100 % · t=+2.1 |
| sports | 2 | 0.578→0.568 -0.010 · B wins 50 % · t=-0.3 | 3.991→3.211 -0.781 · B wins 0 % · t=-2.0 | 4.758→4.973 +0.215 · B wins 50 % · t=+0.2 | 2.120→2.141 +0.021 · B wins 50 % · t=+0.1 | 0.094→0.016 -0.078 · B wins 0 % · t=-1.0 | 0.439→0.497 +0.058 · B wins 100 % · t=+1.2 | 0.600→0.609 +0.009 · B wins 50 % · t=+1.0 | 3.349→3.354 +0.004 · B wins 50 % · t=+0.1 | 4.920→5.720 +0.800 · B wins 50 % · t=+1.0 | 2.861→2.599 -0.261 · B wins 0 % · t=-3.5 | 6.395→5.760 -0.635 · B wins 0 % · t=-1.0 | 2.117→2.292 +0.175 · B wins 50 % · t=+0.9 |
| voicenet | 456 | 0.643→0.643 -0.001 · B wins 50 % · t=-0.2 | 1.116→1.139 +0.023 · B wins 48 % · t=+0.8 | 2.978→2.894 -0.085 · B wins 51 % · t=-0.9 | 0.784→0.858 +0.074 · B wins 55 % · t=+1.9 | 0.099→0.101 +0.001 · B wins 21 % · t=+0.5 | 0.535→0.542 +0.006 · B wins 53 % · t=+1.6 | 0.667→0.662 -0.004 · B wins 44 % · t=-0.5 | 3.388→3.389 +0.002 · B wins 53 % · t=+0.3 | 10.083→10.675 +0.592 · B wins 51 % · t=+3.2 | 1.581→1.627 +0.047 · B wins 51 % · t=+1.3 | 2.362→2.297 -0.065 · B wins 46 % · t=-2.8 | 2.017→1.986 -0.031 · B wins 50 % · t=-0.8 |
Over all 161,600 Profile-A candidates. If best-of-200 is worth its compute, the rank-0 take has to differ from the pool by more than noise.
| metric | all candidates | top-10 mean | rank-0 mean | rank-0 lift |
|---|---|---|---|---|
| genuineness 0–6 | 0.338 | 0.722 | 0.999 | +1.41 σ |
| vocal-burst blend 0–10 | 2.104 | 3.012 | 3.243 | +0.57 σ |
| target strength | 0.362 | 0.715 | 0.791 | +0.20 σ |
| WER | 0.143 | 0.097 | 0.091 | -0.25 σ |
| speaker similarity | 0.475 | 0.515 | 0.515 | +0.22 σ |
| naturalness | 0.458 | 0.537 | 0.639 | +1.28 σ |
| quality | 3.402 | 3.406 | 3.400 | -0.01 σ |
| duration s | 8.769 | 9.281 | 9.554 | +0.18 σ |
Read the WER row in the opposite direction: lower is better, so a negative lift
there is the ranking working. The published post-mortems are blunt that
(1−WER) dominates this composite — that is visible here and is the
reason every component is stored separately.
Selected take, Profile A. Read these against the calibration that produced the thresholds: among 170 pairs a listener confirmed were the same speaker, the median ECAPA cosine was 0.632 and 55 % fell below 0.68. A cell at 0.5 is not a failed clone.
| block | groups | mean spk_sim | median | below floor 0.4 | below resample 0.58 |
|---|---|---|---|---|---|
| edge | 28 | 0.459 | 0.445 | 7.1 % | 82.1 % |
| emotion | 320 | 0.490 | 0.483 | 9.7 % | 69.1 % |
| explicit | 2 | 0.553 | 0.553 | 0.0 % | 50.0 % |
| sports | 2 | 0.439 | 0.439 | 0.0 % | 100.0 % |
| voicenet | 456 | 0.535 | 0.500 | 0.2 % | 67.3 % |
| all | 808 | 0.515 | 0.489 | 4.2 % | 68.6 % |
| count | |
|---|---|
| groups completed | 1616 |
| groups skipped by the safety gate | 0 |
| groups that errored | 0 |
| candidates generated | 323,200 |
empty decodes (audio_codes_list empty) | 323 (0.100 %) |
| groups whose selected take is below the resample threshold 0.58 | 1096 |
Empty decodes are counted rather than assumed away: on an earlier brief
every candidate of a non-verbal group returned an empty
audio_codes_list — reference-conditioned and unconditioned, at every token
budget. A zero here is evidence, not the absence of a check.
char_human scaffold cost
speaker identity?The published contained-emotion recipe puts
char_genuine/human@1.0 under both the free and the contained take. With a
reference-conditioned generation that is a fair worry: a character adapter pushing the
model toward a generic human voice could pull it away from the specific reference. Tested
directly on 16 paired emotion groups (32 candidates each, identical seeds,
captions, sampling and text; the only difference is whether char_human is in the
merge spec).
| metric | with scaffold | without scaffold | Δ | without wins | t |
|---|---|---|---|---|---|
| speaker similarity (selected take) | 0.464 | 0.467 | +0.002 | 56 % | +0.12 |
| speaker similarity (group mean) | 0.415 | 0.436 | +0.021 | 69 % | +1.90 |
| target emotion strength | 0.881 | 1.174 | +0.293 | 56 % | +1.78 |
| genuineness | 0.988 | 1.115 | +0.127 | 62 % | +0.80 |
| burst blend | 2.767 | 2.992 | +0.225 | 62 % | +0.61 |
| WER | 0.092 | 0.125 | +0.033 | 44 % | +1.55 |
| composite score | 0.499 | 0.498 | -0.001 | 56 % | -0.05 |
Conversion is deliberately not run inside candidate selection, so that Profile A and
Profile B differ in exactly one thing. It is measured afterwards on the selected take of every
group that landed below the documented VC threshold of 0.45. Keep rule:
sim_after > sim_before and
dnsmos_after ≥ dnsmos_before − 0.15.
| profile | clips | sim before | sim after | Δ sim | improved | DNSMOS before | DNSMOS after | Δ DNSMOS | kept |
|---|---|---|---|---|---|---|---|---|---|
| A | 271 | 0.381 | 0.623 | +0.242 | 100 % | 3.421 | 3.312 | -0.109 | 12 % |
| B | 246 | 0.381 | 0.623 | +0.242 | 100 % | 3.416 | 3.335 | -0.081 | 11 % |
dnsmos_after ≥ dnsmos_before − 0.15).The practical reading: voice conversion is a rescue for takes that have genuinely lost the speaker, not a routine polish. Both the converted and the original audio are retained, along with both similarity and both DNSMOS values, so the trade can be re-decided per downstream use without re-running anything.
The grid holds the original reference audio plus, for every condition, the rank-0 candidate under Profile A and under Profile B side by side, with the measured numbers under each.