Five merge weights λ = 0, 0.25, 0.50, 0.75, 1.00 over a stratified 10 % subset — 86 condition groups, 32 candidates each, 13,760 generations. Everything else held fixed.
The Mediathek emotion LoRA is trained on real broadcast speech. The synthetic takes are weakest on genuineness (does this sound like a person actually feeling it, rather than performing it) and vocal-burst blend (how naturally breaths, laughs and catches sit inside the line). The hypothesis was that merging more real-speech LoRA would raise both.
Everything except the merge weight is identical across the five arms: same reference voice, same emotion adapters at the same strengths, same corrected cue text, same sentences, same seeds, N=32, batch 64. So any difference is the dose.
| metric | λ=0.00 | λ=0.25 | λ=0.50 | λ=0.75 | λ=1.00 | paired Δ at 1.00 |
|---|---|---|---|---|---|---|
| genuineness | 0.376 | 0.374 | 0.376 | 0.395 | 1.191 | +0.815 t=+22.7 |
| vocal-burst blend | 2.044 | 2.012 | 1.968 | 1.939 | 2.317 | +0.273 t=+1.4 |
| emotion strength | 1.293 | 1.277 | 1.269 | 1.265 | 1.495 | +0.203 t=+5.8 |
| speaker similarity | 0.473 | 0.477 | 0.478 | 0.482 | 0.460 | -0.013 t=-1.8 |
| audio quality (DNSMOS) | 3.402 | 3.403 | 3.404 | 3.401 | 3.416 | +0.014 t=+2.1 |
| word error rate | 0.118 | 0.120 | 0.118 | 0.128 | 0.371 | +0.253 t=+17.2 |
| words / second | 2.752 | 2.751 | 2.755 | 2.738 | 1.732 | -1.020 t=-13.4 |
| clip length (s) | 8.437 | 8.435 | 8.463 | 8.626 | 24.875 | +16.437 t=+18.9 |
| vulnerability | 1.782 | 1.752 | 1.756 | 1.771 | 2.028 | +0.245 t=+5.7 |
| arousal | 2.809 | 2.809 | 2.808 | 2.819 | 2.791 | -0.017 t=-0.3 |
| tension | 2.098 | 2.097 | 2.104 | 2.112 | 2.198 | +0.100 t=+2.5 |
| narration index | 1.043 | 1.036 | 1.031 | 0.992 | -0.650 | -1.692 t=-18.4 |
The paired column compares each group against itself at λ=0, so it removes between-group variance; t beyond about ±2 is reliable. Note how little separates the first four columns.
| metric | Δ at λ=0.25 | Δ at λ=0.50 | Δ at λ=0.75 |
|---|---|---|---|
| genuineness | -0.002 t=-0.3 | -0.000 t=-0.0 | +0.019 t=+1.8 |
| vocal-burst blend | -0.032 t=-1.0 | -0.076 t=-1.3 | -0.105 t=-1.6 |
| emotion strength | -0.016 t=-1.3 | -0.023 t=-1.3 | -0.028 t=-1.3 |
| speaker similarity | +0.004 t=+1.5 | +0.005 t=+1.6 | +0.009 t=+2.4 |
| audio quality (DNSMOS) | +0.001 t=+0.9 | +0.002 t=+1.6 | -0.001 t=-0.5 |
| word error rate | +0.002 t=+1.2 | +0.000 t=+0.2 | +0.010 t=+2.6 |
This confirms and extends the earlier finding that a 25 % merge is nearly a null result. It is not that 25 % is too small a step — 50 % and 75 % do not help either. There is no useful operating point on this curve below the one where the model falls apart.
| λ | median length | 95th pct | at the 32 s cap | WER > 0.5 | slower than 1.5 words/s |
|---|---|---|---|---|---|
| 0.00 | 7.36 s | 14.88 s | 0.3 % | 3.1 % | 4.0 % |
| 0.25 | 7.36 s | 14.79 s | 0.2 % | 3.3 % | 4.1 % |
| 0.50 | 7.36 s | 14.88 s | 0.3 % | 3.1 % | 4.2 % |
| 0.75 | 7.44 s | 15.12 s | 0.8 % | 3.9 % | 5.5 % |
| 1.00 | 32.00 s | 32.00 s | 68.9 % | 30.8 % | 58.9 % |
At every dose up to 0.75 the median take is about 7.4 seconds and roughly one clip in a hundred reaches the cap. At full strength the median take is the cap: median, 95th percentile and maximum are all 32.0 seconds. The model generates until it is cut off, 68.9 % of the time, and 58.9 % of takes fall below 1.5 words per second — drifting, drawn-out speech rather than a delivered line.
| n | genuineness | WER | emotion strength | |
|---|---|---|---|---|
| λ=1.00, longer than 20 s | 1,942 | 1.260 | 0.445 | 1.678 |
| λ=1.00, 20 s or under | 810 | 1.026 | 0.195 | 1.058 |
| λ=0 baseline | 2,752 | 0.376 | 0.118 | 1.293 |
The narration index agrees, and it is the clearest single tell: it falls from 1.043 at λ=0 to -0.650 at full strength — it changes sign. The takes stop reading as someone delivering a line to a listener.
The same group at no merge, three-quarter merge and full merge. The first three are where the genuineness score rose most — which is exactly where you can hear the model losing the line.
Ship λ=0. There is no dose of this LoRA that improves the corpus. Below full strength it does nothing measurable; at full strength it produces runaway, slow, poorly-articulated takes 69 % of the time. The small speaker-similarity gain at 0.25 (+0.004) is real but negligible and does not justify the extra adapter in the merge.
The wider lesson is about the metric, not the adapter. Genuineness moved +0.815 — by far the largest single movement in this sweep — and it was pointing at a broken generator. Any selection rule that weights genuineness heavily would have preferred exactly the takes that stopped saying the sentence. This is a live risk in the ranking formula, not a hypothetical.
For contrast, the intervention that did work on these same dimensions was changing the words rather than the weights: rewriting each line so its content justifies the target emotion raised genuineness 27 % with audio quality flat and word error rate down 12 %. That analysis is here →
max_new_frames=400), not a natural
stopping point. How much longer these takes would run without it is unknown — the true
degeneration may be worse than the duration column suggests.Voice profile k325_age3_bg1 · 86 groups ×
32 candidates × 5 doses = 13,760 generations · 0 failed decodes ·
MOSS-TTS on GH200, batch 64.