MOSS voice profiles
A conditioned-speech corpus generator: for one reference voice, every emotion at four
intensity/containment settings, every VoiceNet dimension at four levels, edge cases and style
variants — in English and German, from the same sentence.
1 · The plan
The full corpus plan & costing →
What gets built, why, and what it costs. 6,000 voices x 832 conditions = 240 M samples,
523k hours. GPU-hours broken out per condition block and per scale (100 / 500 / 1,000 / 6,000
voices), text sourcing, hard-for-TTS domains, vocal-burst balancing, storage. Written for
outsiders.
2 · Listen — speech
Demo grid: old sentence vs written-for-the-condition →
The comparison, side by side in one table. Same voice, same adapters, same seeds — only
the words differ. All 808 groups, both languages, every emotion at all four settings.
Complete voice profiles →
Three full profiles of the same reference voice, all 832 conditions each, with every VoiceNet
dimension at four levels so the extremes sit side by side.
— Profile RW: sentences written per condition
The current best configuration.
— Voice-converted A/B →
Separate Space. All 832 conditions again, each next to its Chatterbox voice-converted
version, level-matched at playback. Conversion buys speaker identity (0.507 → 0.726, every
group) and costs emotional fidelity — both numbers on every row. Best of three noise
draws, selected on the 40-dimension emotion vector.
— Profile A: borrowed sentences
The baseline it is compared against.
— Profile B: +25 % Mediathek LoRA
Kept so the null result is audible.
— Original demo grid
The first grid, 1,617 clips. Superseded by the pages above.
3 · Listen — music
LeVo 2: solo voice, and why the genres missed →
Evolutionary prompt search. The genre miss is an out-of-vocabulary problem — LeVo ships 27
genre tags and "opera", "choir" and "a cappella" are not among them, so it falls back to pop. The
backing choir is a chorus-section artefact. Solo compliance 66 % → 98 %,
genre match 71 % → 96 %, with 115 audio players.
Making LeVo 2 fast enough to scale →
From 0.67× realtime to 11.7× — a 12.7× speedup, turning the pipeline's bottleneck into its
fastest arm. Batching (the code asserted batch 1), static KV cache, CUDA graphs, and a CPU-dispatch
bound decode loop. Includes two batch>1 correctness bugs found in stock LeVo.
How many singing genres are distinguishable? →
153 solo-voice styles, 765 clips, same lyrics throughout. The answer is ~50–80, not 153 —
agreed by three independent representations. Includes a 23040-d music embedding that sees nothing
at all, and whether the model can sing as a child.
Song models — vocals only, with and without stem separation →
Eight arms judged twice: as generated, and as an isolated vocal stem. Three separators
(BS-Roformer, MelBand Roformer, Demucs v4) with measured bulk throughput. Includes the library
default that made the first benchmark meaningless — with it on, a 30-second clip transcribes as
the single word "so". All audio 160 kbps stereo.
4 · Analyses
5 · Reference
What it is, in plain language
A speech model can be asked to say the same sentence angrily, quietly angry,
softly amused, very fast, with a lot of vocal fry, and so on. This project
does that systematically: for a single voice it produces every combination of a large
condition grid, keeps 32 candidates of each (200 in this demo), scores them all, and stores
everything.
The reason for the German-and-English pairing is the eventual goal: a speech-to-speech
translation model that keeps the speaker and the emotional level. To train that you need the
same meaning, spoken by the same voice, at the same emotional intensity, in both languages. That
is what a "group" here is.
The condition matrix — 832 groups per voice
| block | conditions | groups |
| Emotions | 40 emotions × {intense, moderate} × {free-flowing, contained} × {EN, DE} | 320 |
| VoiceNet dimensions | 57 dimensions × 4 levels × {EN, DE} | 456 |
| Edge cases | screams, groans, shivers, laughter, whimpering, crying × {EN, DE} | 28 |
| Sports commentator | × {EN, DE} | 2 |
| Explicitness | × {EN, DE}, subject to an age safety gate | 2 |
| Character clusters | 12 character-cluster LoRAs × {EN, DE}, each also
rendered through Chatterbox voice conversion for comparison | 24 |
| total per voice | 832 |
Every condition uses the measured best practice for that specific emotion or dimension — the
right adapter at the right merge strength with the right prompt phrasing. English and German
share the same text slot in 404 of 404 cells, so the pairing the goal depends on actually
holds.
What was run
| generations | 323,200 per profile |
| candidates stored | 646,917 |
| GPU-hours | 152 |
| throughput | 2,126 generations/GPU-hour |
| failed decodes | 0.10 % |
Every candidate — including the worst-ranked — carries 57 VoiceNet scores, 40 EmoNet scores,
genuineness, vocal-burst blend, a 768-d VoiceCLAP embedding, a 192-d speaker embedding with its
cosine to the reference, Parakeet-v3 word-level timestamps, a procedural caption, and every score
component that produced its rank. Stored as WebDataset tar + parquet, so re-ranking never requires
regenerating anything.
Scaling — and the thing that dominates it
Batch size, not the inference engine, is the lever.
Measured on this exact workload — LoRA merged, reference conditioning, the real prompts:
batch 8 → 3.92× realtime, batch 16 → 5.69×, batch 32 → 7.69×, batch 64 → 8.88×.
The demo ran at batch 8; everything now runs at 64. 2.3× for a one-line change, with no new
engine, no new dependency, and LoRA adapters still working.
Two things this page previously claimed and that turned out to be wrong: that pipelined decode
would give ~3.6× (the port was written and benchmarked — it gives 1.01×, because decode is
not a meaningful share of the loop), and that 6,000 voices at the full matrix "is not feasible"
(at batch 64 it is ~10 % of one budget).
| plan (832 groups/voice) | GPU-h @ batch 64 | % REFORMO left |
| 5,000 × N=16 | 41,897 | 4.4 % |
| 5,000 × N=16 + 900 × N=32 + 100 × N=128 | 63,604 | 6.7 % |
| 5,700 × N=32 + 300 × N=200 | 126,491 | 13.3 % |
| 6,000 × N=32 (flat) | 100,253 | 10.5 % |
| 6,000 × N=128 (max quality) | 400,113 | 41.9 % |
→ Full cost scenarios — ten plans × three throughputs, against
both budgets, with every assumption stated.
Findings
The text matters more than any conditioning change measured so
far. A dozen emotions produced almost no measurable emotion — Bitterness sat at −0.005,
Contempt 0.041, Astonishment 0.039 — even with the correct LoRA loaded and active. The cause was
not the conditioning. It was that the sentences came from a translation corpus, so a group asking
for
intense, freely expressed bitterness was handed a neutral line about a gym membership.
Rewriting each line so its content justifies the target emotion, holding everything else fixed,
raised
emotion strength +22 %, genuineness +27 % and vocal-burst blend +39 %, left audio
quality unchanged and
lowered word error rate 12 %. Bitterness −0.005 → 0.613, Contempt
0.041 → 0.842, Hope/Enthusiasm −0.001 → 1.390, Astonishment 0.039 → 2.347; 24 of 28 conditions
improved. This single change is
larger than every conditioning-side intervention on this pipeline combined, and the text pass costs
~3 GPU-h against ~100,000 for the speech.
Full analysis, with audio →
German conditioning holds the speaker much less well than
English. Measured over all 323,200 candidates: mean speaker similarity 0.578 in English
versus 0.376 in German. German takes are also longer and slower with slightly worse
transcription accuracy. Since the entire point is a German↔English corpus that preserves speaker
identity, this is the central obstacle, not a footnote, and it should be addressed before
scaling.
- Intensity, not containment, breaks the voice clone. Intense conditions sit 0.088 lower
on speaker similarity, with 19.4 % below the usable floor versus 0.0 % for moderate ones.
Holding an emotion back is cheap; pushing it hard is not.
- Containment behaves as designed. The contained conditions shift the vulnerability
dimension by −0.329 while leaving emotional strength essentially intact (−0.022), at a cost of
−0.104 genuineness.
- No dose of the Mediathek LoRA improves the corpus. A full sweep at λ = 0/0.25/0.50/0.75/1.00
over 13,760 generations: quarter, half and three-quarter merges are indistinguishable from no LoRA
on every metric. At full strength genuineness jumps +0.815 — but 68.9 % of takes run to the 32 s
generation cap and word error rate triples, so the model has stopped saying the sentence. The
genuineness scorer rates those runaway takes highest, which is a live risk in the ranking
formula. Ship λ=0. Full sweep, with audio →
- Voice conversion is not free where it is most needed. On 271 badly drifted clips
(below 0.45) it raised similarity 0.381 → 0.623 on every single one, but cost −0.109 DNSMOS,
and only 12 % passed the published keep rule. Earlier figures suggesting conversion is free were
measured on clips that were already close to the target identity.
Honest limitations
- The text rewrite hurts low-arousal conditions. Contemplation, Emotional Numbness, sports
and Interest all lost emotion strength, because the rewrite instruction pushes toward
expressiveness and those targets call for flatness. The instruction needs inverting for them; that
is not yet done.
- No human listening study. Every number here comes from an automatic sensor stack whose
agreement with a human listener on spoken material is only about ρ ≈ +0.21. Treat rankings as
useful, not authoritative.
- The four-level VoiceNet ladder ("extremely low" … "very high") and its 40 % moderate
dose are this project's own construction — the source manual defines seven ordinal levels, not
four — and were not separately validated.
- The A/B/C/D stage directions are a reconstruction; only the four short sublabels are
published.
- Domains are inferred from URL host plus a keyword lexicon (14 domains, ~285–290 pairs
each), because the translation corpus ships a language table rather than a domain taxonomy. The
inferred label is stored per row so it can be replaced.
- The age safety cutoff was recalibrated to 2.40, not the intuitive 3.00: the age
regressor reads systematically low against this dataset's own age labels, so 3.00 would discard
49 % of adults while still admitting 7 % of minors. The gate is conjunctive across three signals.
- In-loop word error rate uses whisper-large-v3-turbo; Parakeet v3 runs as a second pass,
so any re-ranking that prefers it needs no regeneration.
Base model
moss-tts-local-transformer-4.55b-voice-acting-v2
· emotion adapters
· vocal-burst adapters
· Mediathek adapters
· voice-acting manual
· mirror on the HF Space