Does the text carry the emotion?

One voice profile, 84 shared condition groups, 32 candidates each. Everything held fixed except the sentence.

The problem

For roughly a dozen emotions the profile produced almost no measurable emotion at all. Bitterness averaged −0.005 on the emotion head across every candidate. Contempt 0.041. Disappointment 0.028. Astonishment 0.039. This was not a scoring artefact — listening confirmed it. The takes were flat, and the emotion classifier's top label for most of them was Interest, the reading you get from someone neutrally reciting a fact.

The emotion LoRA was loaded and active for every one of those groups; that was verified directly. So the interesting question was where the failure actually lived.

Two candidate explanations.
(a) The conditioning side is broken — the LoRA or the free/contained cue fails to drive these particular emotions.
(b) The words give the actor nothing to feel, so there is nothing to perform.

Explanation (a) was tested first, and rejected. The cue text had a real defect — a single generic “warm and open” free-expression cue was being used for all 40 emotions, including contempt and bitterness, where it is actively wrong. That was fixed to five family-specific cues. It moved nothing: cold emotions scored 0.323 with the free cue and 0.332 with the contained one, a difference indistinguishable from noise. The cue was a genuine bug, but it was not this bug.

That left (b). The sentences came from a translation corpus, so a group asking for intense, freely expressed bitterness was handed a neutral line about a gym membership in Melbourne. No actor can be bitter about that sentence, because the sentence contains nothing to be bitter about.

The test

Each text slot keeps its original translation-corpus line as a seed, and Gemma 12B (bf16) rewrites it so the content gives the speaker a real reason to feel the target emotion — same topic, same register, roughly the same length. Then the whole profile is regenerated with everything else identical: same reference voice, same emotion LoRAs, same corrected cues, same sampling, N=32 candidates per group, batch 64, Mediathek LoRA at λ=0.

Why rewrite instead of writing from scratch? The text pool was deliberately balanced across 14 domains. Writing fresh lines for each emotion would quietly collapse that balance into whatever topic the model finds most emotionally convenient — every bitterness line about a betrayal, every fear line about a dark room. Seeding from the original preserves the domain spread and isolates the variable being tested.
Keeping EN and DE aligned. The corpus depends on the English and German lines of a slot being the same sentence. So the English is rewritten first, and the German is produced as a translation of that rewrite — never rewritten independently, which would silently break the pairing the whole project exists to create.

Result

emotion strength
+22%
1.293 → 1.583
genuineness
+27%
0.376 → 0.478
vocal-burst blend
+39%
2.044 → 2.838

Averaged over every one of the 32 candidates in all 84 shared groups — not cherry-picked takes:

metricoriginal textrewritten textΔrelative
emotion strength1.29261.5828+0.2903+22.5%
genuineness0.37610.4784+0.1023+27.2%
vocal-burst blend2.04382.8384+0.7945+38.9%
audio quality (DNSMOS)3.40203.3944-0.0076-0.2%
word error rate0.11800.1040-0.0140-11.8%
words / second2.75242.8629+0.1105+4.0%
clip length (s)8.43747.6959-0.7415-8.8%

The take that actually ships

Production keeps the best-scoring of the 32 candidates, so the mean over all candidates understates what a listener receives. Best-of-32 per group:

metricoriginal textrewritten textΔ
emotion strength1.76792.2061+0.4383
genuineness1.10881.3861+0.2773
vocal-burst blend4.72075.9063+1.1856
audio quality (DNSMOS)3.49883.4968-0.0020
Read the quality row carefully. Audio quality is flat (3.402 → 3.394, a -0.2% change that is pure noise) and word error rate improved 12%. That matters: it means the emotion was not bought by degrading the audio or by mumbling. The rewrites are simply more speakable than translation-corpus prose — they were written to be said out loud, and the ASR agrees.

Where it came from

The gain is not spread evenly. It lands almost entirely on the emotions that were dead before — which is exactly what explanation (b) predicts:

emotionoriginal textrewritten textΔ
Astonishment Surprise0.0392.347+2.308
Infatuation-0.0151.824+1.839
Disappointment0.0281.499+1.470
Disgust0.0811.484+1.403
Hope Enthusiasm Optimism-0.0011.390+1.391
Embarrassment-0.0101.233+1.243
Helplessness0.0191.123+1.105
Relief0.0601.162+1.101
Amusement0.1851.218+1.033
explicit1.6612.669+1.008
Fear-0.0150.885+0.901
Awe0.7331.570+0.837
Distress0.0000.818+0.818
Contempt0.0410.842+0.801
Bitterness-0.0050.613+0.619
Fatigue Exhaustion0.6731.241+0.568
Confusion0.0090.564+0.555
edge2.8023.206+0.404
Sexual Lust0.2250.533+0.308
Shame0.0060.314+0.307
Elation0.0250.309+0.284
Intoxication Altered States of Consciousness-0.0270.011+0.037
Jealousy and Envy0.0350.069+0.033
voicenet1.8821.903+0.021
Interest2.3352.284-0.051
sports2.0091.764-0.245
Emotional Numbness0.9440.681-0.263
Contemplation0.9780.714-0.264

24 of 28 conditions improved. The emotions that were pinned near zero — Bitterness, Contempt, Disappointment, Astonishment, Embarrassment, Disgust, Fear — are the ones that moved most. Emotions that already scored well moved least, because their seed sentences already happened to contain something to feel.

Four conditions got worse, and they have something in common. Contemplation (−0.264), Emotional Numbness (−0.263), sports (−0.245) and Interest (−0.051) are the low-arousal and neutral-adjacent conditions. The rewrite instruction pushes toward expressiveness, which is the wrong direction when the target is flatness: a numb speaker is supposed to sound like they feel nothing. The instruction needs to be inverted for these — ask for content that justifies detachment, not content that provokes feeling. That is a known, unfixed limitation of this run.

By condition

A = intense & freely expressed, B = intense but contained, C = moderate & free, D = moderate & contained.

conditionemotion strengthgenuinenessvocal-burst blend
origrewrittenorigrewrittenorigrewritten
A0.1901.2000.2890.5562.123.60
B0.2080.6610.1190.4011.502.62
C0.4171.0350.2500.4631.872.71
D0.1690.7920.1730.3391.972.69

Condition A gains most (+1.01). That is the expected shape: when the actor is allowed to express the emotion freely, having something real to express is worth the most. Condition B — intense but held back — gains least (+0.45), because containment caps how much of it can reach the surface anyway.

By language

languageemotion strengthgenuineness
origrewrittenorigrewritten
English1.2711.6400.3880.470
German1.3141.5250.3640.487

Both languages gain. English gains more emotion strength, German more genuineness — consistent with the German side starting from a weaker conditioning baseline, where a clearer text helps the model commit to a reading.

What the rewrites actually look like

Seed on the left, rewrite on the right, with the shipped take from each run. These are the four biggest movers followed by the two worst regressions.

Astonishment Surprise Δ +2.308

original text
The texture of Manouri makes it ideal for crumbling over salads or pasta, or you can simply top the slices with tomatoes, oregano, onions or peppers as an appetizer.
emotion strength of this take: 0.005
rewritten text
I can’t believe how perfectly Manouri crumbles over pasta—but honestly, just topping a slice with tomatoes and oregano is enough to completely blow your mind!
emotion strength of this take: 2.330

Infatuation Δ +1.839

original text
Just give me a moment here, do not touch it, every little movement sends this deep, grinding ache right through the whole bone.
emotion strength of this take: 0.043
rewritten text
Just give me a moment, please don't move, every tiny shift of your body sends this deep, grinding ache right through my very soul.
emotion strength of this take: 2.386

Disappointment Δ +1.470

original text
The goal of the foundation is to prevent malnutrition and undernutrition among young people.
emotion strength of this take: 0.905
rewritten text
We were supposed to stop kids from starving, but we're still failing them.
emotion strength of this take: 2.041

Disgust Δ +1.403

original text
Solution 1: Check Your Malware and Antivirus Software Antivirus programs like Norton and McAfee can take up a large portion of your working memory and cause problems during shutdown.
emotion strength of this take: 0.009
rewritten text
Solution 1: Check your malware and antivirus. These bloated, parasitic programs like Norton and McAfee fester in your system, choking your memory and rotting the shutdown process.
emotion strength of this take: 2.097

Emotional Numbness Δ -0.263

original text
Of course, I need help from time to time, as everyone does – but I don't need it from an eternal bodyguard who doesn't let me solve my own conflicts, but from courageous people who intervene in such situations.
emotion strength of this take: 0.949
rewritten text
Everyone needs help sometimes, I get that. But I don’t need a shadow who suffocates my choices. I just needed someone with enough spine to step in when things got real.
emotion strength of this take: 1.011

Contemplation Δ -0.264

original text
If signs of an allergic reaction occur during therapy, the patient should stop using the drug and consult a doctor.
emotion strength of this take: 0.892
rewritten text
If they start showing signs of an allergic reaction while on the therapy... they have to stop the drug immediately and find a doctor.
emotion strength of this take: 0.887

A safety failure this run produced, and the filter that now blocks it

This has to be written down, because it was caused by the rewrite mechanism itself and it would recur in any similar pipeline.

One seed sentence discussed “the sexualization of society… spread to children”. It happened to be assigned to the Infatuation slot. Asked to rewrite that sentence so the speaker feels infatuation about its topic, the model produced a line savouring “the corruption of society bleeding into the innocence of children”. That is sexualised content about minors, and it was generated from an innocuous seed purely because of the emotion instruction attached to it.

The mechanism: a neutral sentence plus a sexual or romantic emotion instruction is not neutral. The instruction is what makes it harmful. Screening the seeds alone, or the outputs alone, would each have missed it.

Three defences now run, and all three are needed:

  1. Pairing gate. A seed mentioning minors is never paired with a sexual or romantic slot (Infatuation, Sexual Lust, Pleasure/Ecstasy, Affection, Teasing) or with the explicitness block. The seed is swapped for a safe one before any generation happens.
  2. Output gate. Any rewrite that mentions minors and uses sexual or savouring language is rejected and reverted to its seed, and logged.
  3. No silent failures. Both gates print exactly what they caught, so a run that quietly blocks nothing can be distinguished from a run where the filter is broken.

The first version of the output gate did not work. It matched \b(sexual|sexualiz)\b — and that trailing \b cannot match “sexualization”, the exact word that mattered. The filter reported zero hits and looked like it was working. It now uses prefix matching. The rerun with all three defences active produced zero unsafe outputs.

Conclusion

The dead emotions were never a conditioning failure. They were a text failure. Giving the model words that contain something to feel raised emotion strength +22%, genuineness +27% and vocal-burst blend +39%, at no cost in audio quality and with a 12% lower word error rate.

For scale: this single change is larger than every conditioning-side intervention measured on this pipeline so far, including the Mediathek LoRA dose sweep, the cue-family fix, and the recalibrated selection thresholds — combined. It also costs almost nothing. A unique, condition-matched text set for an entire voice profile runs about 3 GPU-hours, against roughly 100,000 GPU-hours for the speech generation it feeds. It is the cheapest large win available.

Two things are still open. The low-arousal conditions need an inverted rewrite instruction, as described above. And the rewrite currently runs once per text slot — whether a second pass conditioned on the measured emotion of the first attempt would help has not been tested.


Voice profile k325_age3_bg1 · 86 shared condition groups · 32 candidates per group per arm · 5,504 generations analysed · Mediathek LoRA λ=0 in both arms · rewriter: Gemma 12B bf16 · generator: MOSS-TTS on GH200.