Mediathek LoRA: how much should you merge?

Five merge weights λ = 0, 0.25, 0.50, 0.75, 1.00 over a stratified 10 % subset — 86 condition groups, 32 candidates each, 13,760 generations. Everything else held fixed.

What was being tested

The Mediathek emotion LoRA is trained on real broadcast speech. The synthetic takes are weakest on genuineness (does this sound like a person actually feeling it, rather than performing it) and vocal-burst blend (how naturally breaths, laughs and catches sit inside the line). The hypothesis was that merging more real-speech LoRA would raise both.

Everything except the merge weight is identical across the five arms: same reference voice, same emotion adapters at the same strengths, same corrected cue text, same sentences, same seeds, N=32, batch 64. So any difference is the dose.

The answer is: don't. 0.25 to 0.75 do nothing, and 1.00 breaks the generator.
Quarter, half and three-quarter doses are statistically indistinguishable from no LoRA at all on every metric. At full strength every number moves sharply — genuineness by +0.815 — but 68.9 % of clips run all the way to the 32 second generation cap and word error rate roughly triples. The model stops saying the sentence. The genuineness gain is a symptom of that, not a benefit.

Every metric at every dose

metricλ=0.00λ=0.25λ=0.50λ=0.75λ=1.00paired Δ at 1.00
genuineness0.3760.3740.3760.3951.191+0.815
t=+22.7
vocal-burst blend2.0442.0121.9681.9392.317+0.273
t=+1.4
emotion strength1.2931.2771.2691.2651.495+0.203
t=+5.8
speaker similarity0.4730.4770.4780.4820.460-0.013
t=-1.8
audio quality (DNSMOS)3.4023.4033.4043.4013.416+0.014
t=+2.1
word error rate0.1180.1200.1180.1280.371+0.253
t=+17.2
words / second2.7522.7512.7552.7381.732-1.020
t=-13.4
clip length (s)8.4378.4358.4638.62624.875+16.437
t=+18.9
vulnerability1.7821.7521.7561.7712.028+0.245
t=+5.7
arousal2.8092.8092.8082.8192.791-0.017
t=-0.3
tension2.0982.0972.1042.1122.198+0.100
t=+2.5
narration index1.0431.0361.0310.992-0.650-1.692
t=-18.4

The paired column compares each group against itself at λ=0, so it removes between-group variance; t beyond about ±2 is reliable. Note how little separates the first four columns.

Doses below full strength are a null result

metricΔ at λ=0.25Δ at λ=0.50Δ at λ=0.75
genuineness-0.002
t=-0.3
-0.000
t=-0.0
+0.019
t=+1.8
vocal-burst blend-0.032
t=-1.0
-0.076
t=-1.3
-0.105
t=-1.6
emotion strength-0.016
t=-1.3
-0.023
t=-1.3
-0.028
t=-1.3
speaker similarity+0.004
t=+1.5
+0.005
t=+1.6
+0.009
t=+2.4
audio quality (DNSMOS)+0.001
t=+0.9
+0.002
t=+1.6
-0.001
t=-0.5
word error rate+0.002
t=+1.2
+0.000
t=+0.2
+0.010
t=+2.6

This confirms and extends the earlier finding that a 25 % merge is nearly a null result. It is not that 25 % is too small a step — 50 % and 75 % do not help either. There is no useful operating point on this curve below the one where the model falls apart.

What actually happens at λ=1.00

λmedian length95th pctat the 32 s capWER > 0.5slower than 1.5 words/s
0.007.36 s14.88 s0.3 %3.1 %4.0 %
0.257.36 s14.79 s0.2 %3.3 %4.1 %
0.507.36 s14.88 s0.3 %3.1 %4.2 %
0.757.44 s15.12 s0.8 %3.9 %5.5 %
1.0032.00 s32.00 s68.9 %30.8 %58.9 %

At every dose up to 0.75 the median take is about 7.4 seconds and roughly one clip in a hundred reaches the cap. At full strength the median take is the cap: median, 95th percentile and maximum are all 32.0 seconds. The model generates until it is cut off, 68.9 % of the time, and 58.9 % of takes fall below 1.5 words per second — drifting, drawn-out speech rather than a delivered line.

Why the genuineness number is misleading. Split the λ=1.00 arm by whether the take ran away:
ngenuineness WERemotion strength
λ=1.00, longer than 20 s1,942 1.2600.445 1.678
λ=1.00, 20 s or under810 1.0260.195 1.058
λ=0 baseline2,7520.376 0.1181.293
The genuineness score is highest exactly where the audio is worst. Even the well-behaved minority at λ=1.00 pays for its genuineness with a higher error rate (0.195 vs 0.118) and lower emotion strength (1.058 vs 1.293). The most defensible reading is that the genuineness scorer rewards breathy, hesitant, vocalisation-heavy audio, and a full-strength merge pushes the model into producing that instead of the line it was asked for. Optimising this metric directly would make the corpus worse.

The narration index agrees, and it is the clearest single tell: it falls from 1.043 at λ=0 to -0.650 at full strength — it changes sign. The takes stop reading as someone delivering a line to a listener.

Listen to it

The same group at no merge, three-quarter merge and full merge. The first three are where the genuineness score rose most — which is exactly where you can hear the model losing the line.

V|HARM|moderately_high|de

“Es ist auch offensichtlich, dass dieses Hormon an den Lernprozessen sowie an der Regulierung der Homöostase beteiligt ist.”
λ = 0.00
7.4 s · genuineness 0.41 · WER 0.00 · 2.42 words/s
λ = 0.75
7.0 s · genuineness 0.10 · WER 0.00 · 2.59 words/s
λ = 1.00
32.0 s (hit the cap) · genuineness 4.40 · WER 0.00 · 0.56 words/s

V|COGL|moderately_low|en

“This is just one of many reasons why it's a good idea to create a social media guide that introduces your employees to social media from a professional standpoint.”
λ = 0.00
7.4 s · genuineness 0.84 · WER 0.00 · 3.90 words/s
λ = 0.75
7.0 s · genuineness 1.06 · WER 0.00 · 4.17 words/s
λ = 1.00
8.5 s · genuineness 2.98 · WER 0.00 · 3.42 words/s

V|R_ORAL|moderately_high|de

“Nach dem ersten Staatsexamen arbeitete sie bei einer internationalen Großkanzlei im Bereich des Business Supports.”
λ = 0.00
8.0 s · genuineness 0.20 · WER 0.00 · 1.88 words/s
λ = 0.75
8.1 s · genuineness 0.05 · WER 0.00 · 1.86 words/s
λ = 1.00
32.0 s (hit the cap) · genuineness 2.67 · WER 0.00 · 0.47 words/s

X|whimpering|en

“Please do not leave me here all alone in the dark like this, I am so scared and I do not know what to do.”
λ = 0.00
12.7 s · genuineness 1.50 · WER 0.12 · 2.04 words/s
λ = 0.75
12.9 s · genuineness 1.91 · WER 0.12 · 2.10 words/s
λ = 1.00
12.9 s · genuineness 1.21 · WER 0.04 · 1.94 words/s

Conclusion

Ship λ=0. There is no dose of this LoRA that improves the corpus. Below full strength it does nothing measurable; at full strength it produces runaway, slow, poorly-articulated takes 69 % of the time. The small speaker-similarity gain at 0.25 (+0.004) is real but negligible and does not justify the extra adapter in the merge.

The wider lesson is about the metric, not the adapter. Genuineness moved +0.815 — by far the largest single movement in this sweep — and it was pointing at a broken generator. Any selection rule that weights genuineness heavily would have preferred exactly the takes that stopped saying the sentence. This is a live risk in the ranking formula, not a hypothetical.

For contrast, the intervention that did work on these same dimensions was changing the words rather than the weights: rewriting each line so its content justifies the target emotion raised genuineness 27 % with audio quality flat and word error rate down 12 %. That analysis is here →

Honest limitations


Voice profile k325_age3_bg1 · 86 groups × 32 candidates × 5 doses = 13,760 generations · 0 failed decodes · MOSS-TTS on GH200, batch 64.