Per-voice conditioned-speech corpus — execution protocol

Every step, adapter, prompt and setting actually used in the two-profile demo run, with the measured result of each.

1. What was actually run2. Measured throughput3. Corpus integrity checks4. Measured results per block5. Profile A vs Profile B: what the 25 % Mediathek merge did5a. Does the ranking select anything?6. Speaker identity across the matrix7. Failures, empties and skips8. Ablation: the char_human scaffold9. Voice-conversion repair, measured10. The demo grid

1. What was actually run

reference voicek325_age3_bg1 — Velvet Sage Baritone, Male, Late 40s to 60s variant chosencbx (best DNSMOS of orig/sidon/cbx) groups808 candidates / group200 profilesA = documented best practice · B = identical + laion/moss-mediathek-emotion-lora :: r64_e2 merged at 25 % total generations323,200 hardware16 nodes × 4 GH200, one worker per GPU, 32 shards × 2 profiles, batch 32 base modellaion/moss-tts-local-transformer-4.55b-voice-acting-v2, bf16, sdpa

Profile B differs from Profile A in exactly one respect: the Mediathek adapter is added to every set_active spec at 0.25. Same seeds, same captions, same sampling, same text, same reference.

2. Measured throughput

blockgroupscandidatesmean gen s/groupmean score s/groupgen / GPU-hourempty decodes
edge5611200329132,1060.00%
emotion640128000331132,0920.01%
explicit4800312122,2200.00%
sports4800157124,2630.00%
voicenet912182400322132,1460.17%
all16163232002,1260.10%

3. Corpus integrity checks

checkresult
EN/DE cells that share a text slot404 / 404 — 0 mismatched
distinct text slots per emotion (must be exactly 1 across its 8 groups)[1]
groups in the matrix808
candidates stored per group200 — every one of them, ranked

A worked example: the four conditions of Fear and both languages all draw the same slot emo::Fear, so the EN and DE takes differ only in language.

ENThe infusion has a mild taste and goes well with herbs, tea, or mate tea.
DEDer Aufguss hat einen milden Geschmack und passt gut zu Kräutern, Tee oder Mate Tee.

4. Measured results per block

Profile A

blockgroupsscorespk_simWERgenuinenessblendstrengthnaturalnessdur sDFLUwords/s
edge280.5230.4590.1551.857.772.920.7328.02.421.89
emotion3200.5920.4900.0740.743.220.600.5929.01.072.70
explicit20.5710.5530.2151.502.522.270.7319.72.341.45
sports20.5780.4390.0943.994.762.120.6004.92.866.39
voicenet4560.6430.5350.0991.122.980.780.66710.11.582.36

Profile B

blockgroupsscorespk_simWERgenuinenessblendstrengthnaturalnessdur sDFLUwords/s
edge280.5150.4600.1531.897.402.800.7667.92.361.90
emotion3200.5910.5010.0760.783.200.570.5919.01.112.71
explicit20.5550.5110.2491.552.821.960.7758.42.171.61
sports20.5680.4970.0163.214.972.140.6095.72.605.76
voicenet4560.6430.5420.1011.142.890.860.66210.71.632.30

The four emotion conditions, Profile A

The comparison the design exists to make. Read alongside the published pooled figures for the same four conditions on six other voices (A emo 0.437 / genu 0.362 / spk 0.569 · B 0.273 / 0.400 / 0.617 · C 0.294 / 0.265 / 0.559 · D 0.185 / 0.326 / 0.617), where the headline was that intensity, not containment, breaks the clone.

conditiongroupsemotion strengthgenuinenessblendspk_simVULNWERnaturalnessbelow floor
A — Intense, free800.7710.9553.5340.4221.9090.0780.64819 / 80
B — Moderate, free800.4800.6223.2230.5371.9720.0730.5800 / 80
C — Intense, contained800.7610.7313.1950.4711.5630.0740.57512 / 80
D — Moderate, contained800.4450.6372.9200.5331.6610.0700.5630 / 80

Containment contrast (C,D − A,B): VULN -0.329, emotion strength -0.022, genuineness -0.104.
Intensity contrast (A,C − B,D): emotion strength +0.304, speaker similarity -0.088, below-floor rate 19.4 % vs 0.0 %.

English vs German, Profile A

The EN/DE pairing is the point of the corpus, so the two halves are reported separately. They use the same sentence, the same voice, the same adapters and the same sampling.

languagegroupsWERspk_simgenuinenessblendemotion strengthdur swords/snaturalness
EN4040.0830.6070.9833.3671.4348.32.790.689
DE4040.1000.4221.0153.1181.47610.82.180.589

5. Profile A vs Profile B: what the 25 % Mediathek merge did

Paired on the rank-0 take of every group that both profiles produced. t is a paired t-statistic on the per-group difference; with hundreds of groups, |t| > 3 is a real effect and |t| < 2 is not.

blocknscoregenuinenessblendstrength_rawwerspk_simnaturalnessqualitydurdfluwpsvuln
ALL8080.619→0.617
-0.001 · B wins 49 % · t=-0.4
0.999→1.030
+0.030 · B wins 48 % · t=+1.5
3.243→3.176
-0.067 · B wins 49 % · t=-1.0
0.791→0.817
+0.027 · B wins 45 % · t=+1.1
0.091→0.093
+0.001 · B wins 21 % · t=+0.7
0.515→0.523
+0.008 · B wins 54 % · t=+2.6
0.639→0.638
-0.001 · B wins 47 % · t=-0.2
3.400→3.401
+0.000 · B wins 53 % · t=+0.0
9.554→9.885
+0.330 · B wins 50 % · t=+2.6
1.414→1.451
+0.037 · B wins 52 % · t=+1.4
2.488→2.454
-0.034 · B wins 46 % · t=-2.1
1.986→1.975
-0.011 · B wins 50 % · t=-0.4
edge280.523→0.515
-0.008 · B wins 36 % · t=-0.8
1.853→1.885
+0.032 · B wins 50 % · t=+0.3
7.771→7.397
-0.373 · B wins 36 % · t=-1.4
2.917→2.797
-0.121 · B wins 32 % · t=-1.1
0.155→0.153
-0.002 · B wins 14 % · t=-0.2
0.459→0.460
+0.001 · B wins 50 % · t=+0.1
0.732→0.766
+0.034 · B wins 50 % · t=+0.8
3.340→3.349
+0.008 · B wins 50 % · t=+0.5
8.000→7.906
-0.094 · B wins 36 % · t=-0.3
2.422→2.364
-0.058 · B wins 50 % · t=-0.2
1.892→1.900
+0.008 · B wins 64 % · t=+0.1
3.852→3.923
+0.071 · B wins 43 % · t=+0.3
emotion3200.592→0.591
-0.001 · B wins 48 % · t=-0.2
0.736→0.782
+0.046 · B wins 49 % · t=+1.5
3.218→3.199
-0.019 · B wins 48 % · t=-0.2
0.596→0.570
-0.026 · B wins 32 % · t=-0.9
0.074→0.076
+0.002 · B wins 21 % · t=+0.7
0.490→0.501
+0.011 · B wins 55 % · t=+2.2
0.592→0.591
-0.001 · B wins 50 % · t=-0.1
3.424→3.421
-0.003 · B wins 53 % · t=-0.7
8.965→8.966
+0.002 · B wins 49 % · t=+0.0
1.074→1.108
+0.034 · B wins 53 % · t=+1.0
2.702→2.710
+0.007 · B wins 45 % · t=+0.3
1.776→1.782
+0.005 · B wins 50 % · t=+0.2
explicit20.571→0.555
-0.016 · B wins 0 % · t=-2.3
1.504→1.546
+0.042 · B wins 50 % · t=+0.1
2.521→2.821
+0.300 · B wins 50 % · t=+0.2
2.267→1.964
-0.303 · B wins 0 % · t=-1.0
0.215→0.249
+0.033 · B wins 50 % · t=+1.0
0.553→0.511
-0.042 · B wins 0 % · t=-1.2
0.731→0.775
+0.043 · B wins 50 % · t=+1.0
3.418→3.459
+0.041 · B wins 100 % · t=+3.0
9.720→8.400
-1.320 · B wins 0 % · t=-4.7
2.345→2.173
-0.171 · B wins 50 % · t=-0.5
1.448→1.614
+0.166 · B wins 100 % · t=+3.3
2.374→2.806
+0.432 · B wins 100 % · t=+2.1
sports20.578→0.568
-0.010 · B wins 50 % · t=-0.3
3.991→3.211
-0.781 · B wins 0 % · t=-2.0
4.758→4.973
+0.215 · B wins 50 % · t=+0.2
2.120→2.141
+0.021 · B wins 50 % · t=+0.1
0.094→0.016
-0.078 · B wins 0 % · t=-1.0
0.439→0.497
+0.058 · B wins 100 % · t=+1.2
0.600→0.609
+0.009 · B wins 50 % · t=+1.0
3.349→3.354
+0.004 · B wins 50 % · t=+0.1
4.920→5.720
+0.800 · B wins 50 % · t=+1.0
2.861→2.599
-0.261 · B wins 0 % · t=-3.5
6.395→5.760
-0.635 · B wins 0 % · t=-1.0
2.117→2.292
+0.175 · B wins 50 % · t=+0.9
voicenet4560.643→0.643
-0.001 · B wins 50 % · t=-0.2
1.116→1.139
+0.023 · B wins 48 % · t=+0.8
2.978→2.894
-0.085 · B wins 51 % · t=-0.9
0.784→0.858
+0.074 · B wins 55 % · t=+1.9
0.099→0.101
+0.001 · B wins 21 % · t=+0.5
0.535→0.542
+0.006 · B wins 53 % · t=+1.6
0.667→0.662
-0.004 · B wins 44 % · t=-0.5
3.388→3.389
+0.002 · B wins 53 % · t=+0.3
10.083→10.675
+0.592 · B wins 51 % · t=+3.2
1.581→1.627
+0.047 · B wins 51 % · t=+1.3
2.362→2.297
-0.065 · B wins 46 % · t=-2.8
2.017→1.986
-0.031 · B wins 50 % · t=-0.8

5a. Does the ranking select anything?

Over all 161,600 Profile-A candidates. If best-of-200 is worth its compute, the rank-0 take has to differ from the pool by more than noise.

metricall candidatestop-10 meanrank-0 meanrank-0 lift
genuineness 0–60.3380.7220.999+1.41 σ
vocal-burst blend 0–102.1043.0123.243+0.57 σ
target strength0.3620.7150.791+0.20 σ
WER0.1430.0970.091-0.25 σ
speaker similarity0.4750.5150.515+0.22 σ
naturalness0.4580.5370.639+1.28 σ
quality3.4023.4063.400-0.01 σ
duration s8.7699.2819.554+0.18 σ

Read the WER row in the opposite direction: lower is better, so a negative lift there is the ranking working. The published post-mortems are blunt that (1−WER) dominates this composite — that is visible here and is the reason every component is stored separately.

6. Speaker identity across the matrix

Selected take, Profile A. Read these against the calibration that produced the thresholds: among 170 pairs a listener confirmed were the same speaker, the median ECAPA cosine was 0.632 and 55 % fell below 0.68. A cell at 0.5 is not a failed clone.

blockgroupsmean spk_simmedianbelow floor 0.4below resample 0.58
edge280.4590.4457.1 %82.1 %
emotion3200.4900.4839.7 %69.1 %
explicit20.5530.5530.0 %50.0 %
sports20.4390.4390.0 %100.0 %
voicenet4560.5350.5000.2 %67.3 %
all8080.5150.4894.2 %68.6 %

7. Failures, empties and skips

count
groups completed1616
groups skipped by the safety gate0
groups that errored0
candidates generated323,200
empty decodes (audio_codes_list empty)323 (0.100 %)
groups whose selected take is below the resample threshold 0.581096

Empty decodes are counted rather than assumed away: on an earlier brief every candidate of a non-verbal group returned an empty audio_codes_list — reference-conditioned and unconditioned, at every token budget. A zero here is evidence, not the absence of a check.

8. Ablation: does the char_human scaffold cost speaker identity?

The published contained-emotion recipe puts char_genuine/human@1.0 under both the free and the contained take. With a reference-conditioned generation that is a fair worry: a character adapter pushing the model toward a generic human voice could pull it away from the specific reference. Tested directly on 16 paired emotion groups (32 candidates each, identical seeds, captions, sampling and text; the only difference is whether char_human is in the merge spec).

metricwith scaffoldwithout scaffoldΔwithout winst
speaker similarity (selected take)0.4640.467+0.00256 %+0.12
speaker similarity (group mean)0.4150.436+0.02169 %+1.90
target emotion strength0.8811.174+0.29356 %+1.78
genuineness0.9881.115+0.12762 %+0.80
burst blend2.7672.992+0.22562 %+0.61
WER0.0920.125+0.03344 %+1.55
composite score0.4990.498-0.00156 %-0.05
Answer: no, it does not cost identity. Speaker similarity on the selected take is flat (Δ +0.002, |t| < 1), which is the question that was asked and the one this answers cleanly. Everything else here is a trend, not a result. Removing the scaffold moves emotion strength (+0.293) and WER (+0.033) in the direction the recipe predicts — it trades expressiveness for intelligibility — but at n = 16 paired groups none of those |t| reach 2, and the composite score is unchanged (-0.001). The scaffold is kept because the published recipe prescribes it and this test found no reason to drop it — not because this test endorsed it. Re-run at corpus scale before treating any row but the first as settled.

9. Voice-conversion repair, measured

Conversion is deliberately not run inside candidate selection, so that Profile A and Profile B differ in exactly one thing. It is measured afterwards on the selected take of every group that landed below the documented VC threshold of 0.45. Keep rule: sim_after > sim_before and dnsmos_after ≥ dnsmos_before − 0.15.

profileclipssim beforesim afterΔ simimprovedDNSMOS beforeDNSMOS afterΔ DNSMOSkept
A2710.3810.623+0.242100 %3.4213.312-0.10912 %
B2460.3810.623+0.242100 %3.4163.335-0.08111 %
This does not replicate “nearly free”, and the reason is selection. The published figure — 68 parts, similarity 0.755 → 0.820, DNSMOS 3.30 → 3.35 — was measured on parts that were already close to the reference. Applied instead to the worst-drifted takes in this corpus (mean similarity 0.38), conversion buys far more identity (+0.242, improving 100 % of clips) but does not come free: DNSMOS falls -0.109 on average, and only 12 % of conversions clear the published keep rule (dnsmos_after ≥ dnsmos_before − 0.15).

The practical reading: voice conversion is a rescue for takes that have genuinely lost the speaker, not a routine polish. Both the converted and the original audio are retained, along with both similarity and both DNSMOS values, so the trade can be re-decided per downstream use without re-running anything.

10. The demo grid

The grid holds the original reference audio plus, for every condition, the rank-0 candidate under Profile A and under Profile B side by side, with the measured numbers under each.

→ open the demo grid (audio)