Measured constants, three inference engines, eight build scenarios, against both compute budgets. Profile B (the Mediathek merge) is dropped — standard profile only.
| batch | × realtime | ms / clip |
|---|---|---|
| 8 (what the demo ran) | 3.92 | 1254 |
| 16 | 5.69 | 859 |
| 32 | 7.69 | 637 |
| 64 (now in use) | 8.88 | 543 |
| block | conditions | groups |
|---|---|---|
| Emotions | 40 × {intense, moderate} × {free, contained} × {EN, DE} | 320 |
| VoiceNet dimensions | 57 × 4 levels × {EN, DE} | 456 |
| Edge cases | screams, groans, shivers, laughter, whimpering, crying × {EN, DE} | 28 |
| Sports commentator | × {EN, DE} | 2 |
| Explicitness | × {EN, DE}, behind the age safety gate | 2 |
| Character clusters NEW | 12 character-cluster LoRAs × {EN, DE} — each also rendered a second time through Chatterbox voice conversion, so the two can be compared later | 24 |
| total per voice | 832 | |
Share of the achievable best-of-200 gain, from the exact order-statistic analysis:
| N | reward | speaker sim | genuineness | emotion |
|---|---|---|---|---|
| 8 | 50% | 64% | 36% | 38% |
| 16 | 63% | 74% | 50% | 50% |
| 32 | 75% | 83% | 63% | 64% |
| 64 | 86% | 90% | 77% | 77% |
| 128 | 95% | 96% | 91% | 91% |
| 200 | 100% | 100% | 100% | 100% |
Speaker similarity saturates first — 74 % of the achievable gain already at N=16, 83 % at 32. Genuineness and emotion strength keep climbing all the way to 200, so they are what larger N actually buys.
a + b·√(ln N) to the
measured points (R² = 0.999 for reward, 0.997 for speaker similarity) gives reward 0.6199 → 0.6957
and speaker similarity 0.6540 → 0.6785 going from N=200 to N=1024 — a 5× cost increase for
+0.076 and +0.024. S10 is included only to show that; S9 (300 × N=200) captures most of it
for half the price and is the one to prefer. Note also that in S9 the 300 high-N voices are
already 26 % of all clips despite being 5 % of the voices; at N=1024 they are 63 %.| project | total core-h | remaining | % left | expires |
|---|---|---|---|---|
| REFORMO | 201,600,000 | 68,709,721 | 34.1% | 2026-12-31 |
| LAIONIZE | 41,660,000 | 37,569,231 | 90.2% | 2027-04-30 |
| scenario | engine | GPU-h | core-h | % REFORMO left | % LAIONIZE left |
|---|---|---|---|---|---|
| S1 6,000 × N=32 (flat) 159,744,000 clips · 391,817 h audio reward: N=32: 75% | batch 8 — as the demo ran (3.9x) | 100,253 | 7,218,230 | 10.5% | 19.2% |
| batch 64 — measured, now in use (8.9x) | 44,423 | 3,198,491 | 4.7% | 8.5% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 12,858 | 925,792 | 1.3% | 2.5% | |
| S2 5,000 × N=32 + 1,000 × N=64 186,368,000 clips · 457,119 h audio reward: N=32: 75% / N=64: 86% | batch 8 — as the demo ran (3.9x) | 116,912 | 8,417,669 | 12.3% | 22.4% |
| batch 64 — measured, now in use (8.9x) | 51,777 | 3,727,973 | 5.4% | 9.9% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 14,951 | 1,076,491 | 1.6% | 2.9% | |
| S3 5,000 × N=16 only 66,560,000 clips · 163,257 h audio reward: N=16: 63% | batch 8 — as the demo ran (3.9x) | 41,897 | 3,016,596 | 4.4% | 8.0% |
| batch 64 — measured, now in use (8.9x) | 18,635 | 1,341,705 | 2.0% | 3.6% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 5,483 | 394,747 | 0.6% | 1.1% | |
| S4 5,000 × N=16 + 1,000 × N=64 119,808,000 clips · 293,862 h audio reward: N=16: 63% / N=64: 86% | batch 8 — as the demo ran (3.9x) | 75,265 | 5,419,073 | 7.9% | 14.4% |
| batch 64 — measured, now in use (8.9x) | 33,393 | 2,404,268 | 3.5% | 6.4% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 9,719 | 699,744 | 1.0% | 1.9% | |
| S5 5,000 × N=16 + 900 × N=32 + 100 × N=128 101,171,200 clips · 248,150 h audio reward: N=16: 63% / N=32: 75% / N=128: 95% | batch 8 — as the demo ran (3.9x) | 63,604 | 4,579,466 | 6.7% | 12.2% |
| batch 64 — measured, now in use (8.9x) | 28,245 | 2,033,631 | 3.0% | 5.4% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 8,254 | 594,255 | 0.9% | 1.6% | |
| S6 5,000 × N=16 + 900 × N=32 + 100 × N=250 111,321,600 clips · 273,047 h audio reward: N=16: 63% / N=32: 75% / N=250: 100% | batch 8 — as the demo ran (3.9x) | 69,955 | 5,036,752 | 7.3% | 13.4% |
| batch 64 — measured, now in use (8.9x) | 31,049 | 2,235,496 | 3.3% | 6.0% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 9,052 | 651,709 | 0.9% | 1.7% | |
| S7 5,000 × N=16 + 900 × N=32 + 100 × N=500 132,121,600 clips · 324,065 h audio reward: N=16: 63% / N=32: 75% / N=500: 100% | batch 8 — as the demo ran (3.9x) | 82,970 | 5,973,813 | 8.7% | 15.9% |
| batch 64 — measured, now in use (8.9x) | 36,794 | 2,649,153 | 3.9% | 7.1% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 10,687 | 769,442 | 1.1% | 2.0% | |
| S8 6,000 × N=128 (max quality) 638,976,000 clips · 1,567,266 h audio reward: N=128: 95% | batch 8 — as the demo ran (3.9x) | 400,113 | 28,808,121 | 41.9% | 76.7% |
| batch 64 — measured, now in use (8.9x) | 176,794 | 12,729,163 | 18.5% | 33.9% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 50,533 | 3,638,368 | 5.3% | 9.7% | |
| S9 5,700 × N=32 + 300 × N=200 ← high-quality core 201,676,800 clips · 494,668 h audio reward: N=32: 75% / N=200: 100% | batch 8 — as the demo ran (3.9x) | 126,491 | 9,107,346 | 13.3% | 24.2% |
| batch 64 — measured, now in use (8.9x) | 56,006 | 4,032,425 | 5.9% | 10.7% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 16,155 | 1,163,142 | 1.7% | 3.1% | |
| S10 5,700 × N=32 + 300 × N=1024 (for comparison) 407,347,200 clips · 999,132 h audio reward: N=32: 75% / N=1024: 100% | batch 8 — as the demo ran (3.9x) | 255,181 | 18,373,007 | 26.7% | 48.9% |
| batch 64 — measured, now in use (8.9x) | 112,815 | 8,122,672 | 11.8% | 21.6% | |
| SGLang-Omni 31.2x — LoRA support UNVERIFIED | 32,323 | 2,327,290 | 3.4% | 6.2% |