← overview
Voice profiles — listen
One reference voice, rendered through all 832 conditions, three different ways.
Every clip below is the take that won its group, with the scores that made it win.
What a "voice profile" is. Take one reference speaker. Ask the model to
say something as that speaker in every combination of a large grid: 40 emotions × {intense,
moderate} × {freely expressed, held in}, 57 voice-quality dimensions × 4 levels, plus
screams, laughter, crying, a sports commentator and more — each in English and German,
from the same sentence. That grid is 832 groups. A profile is all of them for one voice.
Profile A — translation-corpus sentences →
The baseline. Sentences drawn from the FineTranslations corpus and used as-is. No Mediathek
LoRA. This is what the pipeline produced before the text-rewrite experiment.
Profile B — the same, plus 25 % Mediathek LoRA →
Identical sentences and conditions, with a real-speech adapter merged at λ=0.25. A
later dose sweep found this to be a null result — kept
published so you can hear that for yourself rather than take it on trust.
Profile RW — sentences rewritten to carry the emotion →
Same voice, same 832 conditions, but every sentence rewritten so its content gives the
speaker a reason to feel the target emotion. Emotion strength +22 %, genuineness +27 %, blend
+39 %, audio quality flat, word error rate down 12 %.
The comparison worth making
Open A and RW side
by side and listen to Bitterness, Contempt or Disappointment in the
"intense · free" column. Same voice, same adapter, same settings — only the words
differ.
| emotion | A (corpus sentences) | RW (rewritten) |
| Bitterness | −0.005 | 0.613 |
| Contempt | 0.041 | 0.842 |
| Disappointment | 0.028 | 1.499 |
| Hope / Enthusiasm | −0.001 | 1.390 |
| Astonishment / Surprise | 0.039 | 2.347 |
Emotion-head score on the winning take. Full analysis:
Does the text carry the emotion?
Audio quality
Profile RW is now 48 kHz at 160 kbps — regenerated at the model's native
rate. Profiles A and B remain 16 kHz mono at ~97 kbps: the pipeline used to resample to 16 kHz
before storing, band-limiting everything to 8 kHz, and that ceiling is baked into their data.
Because of that,
do not A/B profile RW against profile A on this page to judge the
text rewrite — you would be hearing the bandwidth as much as the words. The
side-by-side demo grid deliberately keeps
both columns at
16 kHz so the only difference there is the text.