An evolutionary prompt search over LeVo 2 / SongGeneration-v2-large. Two independent failures — "opera" coming out as auto-tuned pop, and a backing choir arriving in clips that were supposed to be one person alone — located, explained and fixed by prompting. 2523 clips judged by Gemini 3.6 Flash on three axes. All audio 160 kbps stereo.
1. The genre miss is an out-of-vocabulary problem. LeVo's
descriptions field feeds the type_info conditioner, and the model
ships its actual tag vocabulary in sample/description/: 2 genders,
27 genres, 8 emotions, 7 timbres. opera is not in it.
Neither is choir, gospel, hymn or
a cappella. Out-of-vocabulary words are not rejected — they are
conditioned on weakly, and the model falls back to its prior, which is pop.
| description given to LeVo | opera G | choir G | pop G | country G | reading |
|---|---|---|---|---|---|
| (empty string) | 1.67 | 2.17 | 4.90 | 4.70 | With no conditioning at all, LeVo sings pop. That is why removing the prompt costs opera and choir everything and costs pop and country nothing. |
female, opera — the out-of-vocabulary word | 3.17 | — | — | — | The word a human would write is the worst opera prompt in the search. |
female, classical — the in-vocabulary anchor | 4.83 | 4.17 | — | — | Two words, both in the tag list. Best opera prompt found. |
[chorus] is a learned cue for stacked, doubled, harmonised vocals, because that is what
choruses are in the training corpus. Delete the chorus and the choir goes with it.descriptions = "<gender>, <IN-VOCABULARY genre tag>, <free style words>" lyrics = "[verse] Line one. Line two. Line three. Line four. Line five. Line six." gen_type = "vocal" temperature = 0.9 top_k = 50 top_p = 0.0 cfg_coef = 2.0
classical carries opera; opera alone does not.[verse] block. No [chorus]. If you need two sections, tag the
second one [verse] as well.[intro-*], [inst-*] or [outro-*]. They are
instrumental by definition and they put instruments back into a vocals-only render.a cappella. Both cost score. The solo is
bought by structure and by gen_type='vocal', not by asking.Baseline is the prompting the pipeline uses today. Every cell is the same lyrics and the same seeds; solo % is the fraction of takes the judge scored 4–5 for "exactly one vocalist", which is the number that matters because the failure is all-or-nothing rather than gradual.
| genre | arm | n | S one voice | G genre | I no instr. | solo % | clean % | genre % | prompt / structure |
|---|---|---|---|---|---|---|---|---|---|
| opera | baseline | 60 | 4.22 | 3.70 | 5.00 | 73% | 73% | 60% | female, classical, opera, emotional, a cappella, no instrumentsvc · cfg 1.5 · top_k 50 |
| evolved | 36 | 4.75 | 4.44 | 5.00 | 92% | 92% | 86% | female, classicalv_only · cfg 1.5 · top_k 250 | |
| choir | baseline | 60 | 3.43 | 3.33 | 4.98 | 52% | 52% | 42% | female, gospel, choir, uplifting, a cappella, no instrumentsvc · cfg 1.5 · top_k 50 |
| evolved | 24 | 4.96 | 4.88 | 5.00 | 100% | 100% | 96% | female, classical, hymn, sacred, emotional, a cappella, solo singer, single voicev_only · cfg 2.0 · top_k 50 | |
| pop | baseline | 57 | 4.35 | 5.00 | 4.84 | 74% | 68% | 100% | female, pop, emotional, a cappella, no instrumentsvc · cfg 1.5 · top_k 50 |
| evolved | 12 | 5.00 | 5.00 | 4.75 | 100% | 92% | 100% | female, pop, pop rock, emotionalvc · cfg 1.5 · top_k 50 | |
| country | baseline | 60 | 3.98 | 4.33 | 4.80 | 65% | 62% | 82% | male, country, emotional, a cappella, no instrumentsvc · cfg 1.5 · top_k 50 |
| evolved | 30 | 5.00 | 5.00 | 5.00 | 100% | 100% | 100% | male, country, solo vocalv_only · cfg 1.5 · top_k 50 |
classical anchor, the sacred style word, the single
[verse] and CFG 2.0. Note also that choir itself must never appear:
what is wanted is choral style sung by one person, and the baseline's
gospel, choir is why it was the worst of the four.Per genre: the baseline take the judge scored lowest for "one vocalist", a typical baseline take, and the best take from the evolved prompt. What the judge wrote is printed next to each.
| genre | arm | audio | S | G | I | what the judge heard |
|---|---|---|---|---|---|---|
| opera | baseline — worst | 1.00 | 3.00 | 5.00 | A female soprano sings a classical operatic melody alone for the first half, before a full classical choir joins in for the chorus, with no instruments present. | |
| baseline — best | 5.00 | 5.00 | 5.00 | I hear a single female classical opera singer performing the requested lyrics a cappella with natural room reverb and no instrumental accompaniment. | ||
| evolved — best | 5.00 | 5.00 | 5.00 | I hear a single operatic soprano voice singing classical lyrics with traditional bel canto technique and vibrato, completely unaccompanied. | ||
| evolved — worst | 2.00 | 5.00 | 5.00 | A trained classical female soprano sings the aria completely a cappella until the last three seconds, where a full classical choir suddenly enters singing 'Silent hour descend'. There are no instrumental sounds present at all. | ||
| choir | baseline — worst | 1.00 | 2.00 | 5.00 | The clip begins with a solo female voice, but at 0:08 a choir/backing vocal ensemble enters to provide full gospel harmonies behind the lead singer; there are no musical instruments present. | |
| baseline — best | 5.00 | 5.00 | 5.00 | A single female voice singing a modal liturgical hymn unaccompanied, with cathedral reverb and no additional vocal layers or instruments present. | ||
| evolved — best | 5.00 | 5.00 | 5.00 | A single female voice sings a sacred liturgical hymn completely unaccompanied, with no additional voices or instruments present. | ||
| evolved — worst | 4.00 | 5.00 | 5.00 | A solo female soprano sings the sacred hymn completely unaccompanied with cathedral reverb, followed by a brief layered vocal phrase at the very end. | ||
| pop | baseline — worst | 2.00 | 5.00 | 3.00 | I hear a female contemporary pop voice with rhythmic finger snaps in the verse, joined by stacked backing vocal harmonies during the chorus. | |
| baseline — best | 5.00 | 5.00 | 5.00 | I hear a single female voice singing a contemporary pop melody completely solo with no backing vocals or instrumental accompaniment. | ||
| evolved — best | 5.00 | 5.00 | 5.00 | A single female voice sings the prompt's lyrics acapella with a contemporary pop delivery, featuring no extra vocal layers or instruments. | ||
| evolved — worst | 5.00 | 5.00 | 2.00 | A single male voice singing modern pop with rhythmic finger snaps in the background, but no other vocal parts or melodic instruments. | ||
| country | baseline — worst | 1.00 | 2.00 | 5.00 | I hear a male lead vocalist singing in a pop/R&B style, joined in the chorus by a stacked vocal choir performing rhythmic 'oh-oh' vocal pads and harmonies, with no musical instruments present. | |
| baseline — best | 5.00 | 5.00 | 5.00 | A single male country voice singing a cappella with natural vocal reverb and no instruments or second voices present. | ||
| evolved — best | 5.00 | 5.00 | 5.00 | A solo male country voice singing the lyrics acapella with no backing vocals or instruments present. |
Generation 2 swept the section scaffold with everything else held fixed (n = 6 per cell). Score shown is S — exactly one vocalist.
| structure | opera | choir | pop | country | verdict |
|---|---|---|---|---|---|
[verse] only, all lines in one block | 4.83 | 5.00 | 4.50 | 5.00 | best |
[verse] ; [verse] (chorus retagged) | 5.00 | 5.00 | 4.67 | 4.67 | as good |
[verse] ; [chorus] (the obvious thing) | 4.33 | 4.33 | 4.17 | 4.83 | baseline |
[silence] ; [verse] ; [chorus] | 5.00 | 3.67 | 5.00 | 5.00 | erratic |
[verse] ; [bridge] | 2.67 | 5.00 | 4.67 | 3.33 | erratic |
[intro-short] ; … ; [outro-short] | 4.50 | 3.33 | 4.17 | 3.67 | hurts |
[verse] ; [inst-short] ; [chorus] | 2.00 | 4.67 | 4.17 | 4.50 | worst |
[inst-short] is the worst single intervention measured — it takes opera from S 4.33 to 2.00. An [inst-*] section is instrumental by definition; asking for a vocals-only render and then writing an instrumental section into the lyrics is a contradiction, and the model resolves it by adding instruments and voices back.Each row adds one phrase to the same base prompt, same seeds, everything else identical. The two rows marked manipulation check add the unwanted thing positively: if those make it worse, the model demonstrably responds to the phrase, and a negation that fails to help is evidence about how negation is handled — not evidence that the words are ignored.
| phrase added | opera S | choir S | pop S | country S | mean S | mean G |
|---|---|---|---|---|---|---|
(base) | 5.00 | 5.00 | 4.17 | 5.00 | 4.79 | 4.71 |
+ 'a cappella' | 3.67 | 5.00 | 3.83 | 5.00 | 4.38 | 4.12 |
+ 'unaccompanied' | 5.00 | 5.00 | 3.67 | 5.00 | 4.67 | 4.92 |
+ 'solo voice' | 5.00 | 4.83 | 3.83 | 5.00 | 4.67 | 4.71 |
+ 'no instruments' | 5.00 | 5.00 | 4.50 | 5.00 | 4.88 | 4.88 |
+ 'piano' [manipulation check] | 5.00 | 5.00 | 4.33 | 4.67 | 4.75 | 4.79 |
+ 'a cappella, solo voice, unaccompanied' | 5.00 | 5.00 | 4.50 | 4.50 | 4.75 | 4.54 |
BASELINE (control) | 3.33 | 4.00 | 4.17 | 4.17 | 3.92 | 4.17 |
| phrase added | opera S | choir S | pop S | country S | mean S | mean G |
|---|---|---|---|---|---|---|
(base) | 4.83 | 5.00 | 5.00 | 5.00 | 4.96 | 4.75 |
+ 'backing vocals' [manipulation check] | 3.17 | 5.00 | 4.00 | 5.00 | 4.29 | 4.38 |
+ 'choir' [manipulation check] | 5.00 | 5.00 | 2.83 | 5.00 | 4.46 | 4.50 |
+ 'no backing vocals' | 5.00 | 4.50 | 2.83 | 5.00 | 4.33 | 4.71 |
+ 'no choir' | 5.00 | 5.00 | 3.00 | 5.00 | 4.50 | 4.54 |
+ 'no backing vocals, no choir, no harmony, no doubling' | 4.67 | 5.00 | 3.67 | 5.00 | 4.58 | 4.75 |
+ 'solo singer, single voice' (positive) | 5.00 | 5.00 | 3.00 | 5.00 | 4.50 | 4.75 |
BASELINE (control) | 3.50 | 4.00 | 4.17 | 4.17 | 3.96 | 4.25 |
female, opera was the worst opera
prompt in the search (S 3.00 / G 3.17) — worse than saying nothing about the genre
beyond classical.a cappella — the single most destructive phrase found. It takes opera's
genre score from 4.17 to 1.83. Almost certainly because "a cappella" labels a cappella
groups in a music corpus: close-harmony vocal bands, the opposite of a solo aria.
unaccompanied is the safe alternative and was the best opera cell in that probe.gen_type='vocal'. Proved by
manipulation check: adding piano did not add a piano — I stayed at 5.00,
because the bgm codebook is overwritten before rendering. The slot cannot add instruments, so it
cannot remove them either; it only perturbs the voice.folk drags them toward generic Western folk. There is no in-vocabulary tag for any of
them — a model capability limit, not a prompting problem.[verse] that buys solo
compliance elsewhere appears to cost rhythmically-driven genres their structure.The 50 medoids of a Ward clustering of the 153-genre taxonomy. Both arms get identical lyrics and identical seeds; base is the natural-language a-cappella caption the pipeline uses today, evo is the recipe above instantiated per genre. solo % = takes scored 4–5 on one-vocalist and 4–5 on no-instrumentation. PASS = solo % ≥ 83 and mean genre ≥ 4.
| family | n | base clean% | evo clean% | base G | evo G | verdict |
|---|---|---|---|---|---|---|
| blues | 2 | 58% | 75% | 2.75 | 3.50 | solo transfers, genre does not |
| theatre | 2 | 42% | 67% | 2.50 | 2.50 | solo transfers, genre does not |
| folk | 3 | 39% | 67% | 3.78 | 3.72 | solo transfers, genre does not |
| metal | 4 | 29% | 62% | 2.79 | 3.67 | solo transfers, genre does not |
| classical | 7 | 50% | 60% | 2.19 | 3.10 | solo transfers, genre does not |
| functional | 5 | 23% | 59% | 2.30 | 2.55 | solo transfers, genre does not |
| world | 11 | 50% | 52% | 1.79 | 1.53 | does not transfer — genre got worse |
| rock | 5 | 53% | 50% | 3.75 | 4.73 | flat |
| soul | 2 | 47% | 50% | 4.62 | 4.75 | flat |
| pop | 5 | 20% | 41% | 4.17 | 4.75 | transfers |
| urban | 4 | 62% | 38% | 3.38 | 3.88 | regresses |
| genre | family | base S/G/I | base solo% | evo S/G/I | evo solo% | Δ | verdict | baseline take | evolved take | cluster it represents |
|---|---|---|---|---|---|---|---|---|---|---|
| Sámi joik sami joik, cyclical wordless vocable chant, guttural and hypnotic | world | 3.3/1.2/5.0 | 33% | 4.5/1.0/4.5 | 67% | +33 | FAIL | 8 genres: elderly_folk, georgian_solo, hawaiian_falsetto, opera_basso, sardinian_cantu, sean_nos, sertanejo | ||
| Andean huayno andean huayno, very high piercing nasal quechua-inflected tone | world | 3.7/3.5/5.0 | 50% | 3.5/1.2/4.0 | 33% | -17 | FAIL | 7 genres: ambient_wordless, min_yo, neo_soul, opera_sop_coloratura, pansori, whisper_sing | ||
| Bedroom pop bedroom pop, whispered soft double-tracked intimacy, unforced small voice | pop | 2.3/4.0/4.5 | 17% | 3.2/4.4/4.2 | 20% | +3 | solo fail | 7 genres: bossa_nova, gamelan_sindhen, lullaby, lullaby_grandmother, raga_bhajan, vocal_fry_sing | ||
| Keening / lament ritual keening lament, sobbing wailing cries, grief-stricken and unmetred | functional | 2.7/1.0/4.7 | 0% | 4.2/1.3/4.7 | 67% | +67 | FAIL | 6 genres: baroque_ornament, carnatic, enka, rnb_contemporary, trance_vocal | ||
| Bulgarian diaphonic bulgarian village singing, bright nasal forward tone, tight ornament shake | world | 4.2/1.7/5.0 | 67% | 3.5/1.0/4.3 | 50% | -17 | FAIL | 6 genres: country_pop, dnb_vocal, gospel_lead, inuit_katajjaq, schlager | ||
| Oratorio / sacred solo oratorio soloist, clean sustained sacred line, minimal vibrato, devotional | classical | 3.0/2.5/5.0 | 33% | 4.2/3.2/4.5 | 50% | +17 | FAIL | 6 genres: byzantine, canzone, klezmer_krekhts, opera_tenor_lirico, peking_opera | ||
| Teen pop teen pop, bright youthful thin-bodied voice, eager and unweathered | pop | 3.2/3.8/5.0 | 33% | 2.8/4.7/5.0 | 33% | 0 | solo fail | 6 genres: citypop, cumbia, hyperpop, kpop, musical_belt | ||
| Field holler unaccompanied field holler, long falling cries, free rhythm, raw and open- | blues | 4.0/3.7/5.0 | 67% | 4.5/2.5/5.0 | 83% | +17 | genre fail | 6 genres: flamenco_cante, jazz_crooner, outlaw_country, soul_classic, uk_drill_sung | ||
| Samba samba, buoyant percussive syncopated phrasing, open joyful tone | world | 3.5/2.0/4.5 | 33% | 2.5/2.3/5.0 | 17% | -17 | FAIL | 6 genres: gospel_bluegrass, gregorian, shape_note, synthpop_80s, throat_overtone | ||
| Cantorial cantorial chazzanut, virtuosic ornamented liturgical cries, huge emotive r | world | 4.0/3.2/5.0 | 67% | 4.5/2.8/5.0 | 83% | +17 | genre fail | 5 genres: arabic_mawwal, persian_avaz, spiritual, synthwave | ||
| Mezzo-soprano mezzo-soprano, warm dark middle register, rich operatic legato | classical | 3.0/2.0/5.0 | 33% | 4.5/3.2/5.0 | 83% | +50 | genre fail | 5 genres: balkan_open, fado, kulning, musical_legit | ||
| Sprechgesang sprechgesang, half-spoken half-sung expressionist delivery, pitched speech | classical | 4.5/2.7/5.0 | 83% | 4.2/3.0/5.0 | 67% | -17 | FAIL | 5 genres: countertenor, humming, rebetiko, vaporwave_croon | ||
| Beatboxing beatboxing, percussive mouth drum sounds, rhythmic and unpitched | functional | 2.5/1.2/4.5 | 0% | 3.2/1.0/4.0 | 17% | +17 | FAIL | 4 genres: melodic_rap, vocalese, whistling | ||
| Maximum belt maximum-effort belting, full-throated sustained loud chest voice, strained | functional | 3.5/3.2/4.5 | 50% | 4.6/3.6/4.4 | 60% | +10 | FAIL | 4 genres: europop_dance, house_diva, pop_ballad | ||
| Punk punk, shouted snarling barked delivery, aggressive and untrained | rock | 3.2/3.7/5.0 | 33% | 3.7/4.8/5.0 | 33% | 0 | solo fail | 4 genres: grime_hook, hardcore_scream, trap_sung | ||
| Clean metal baritone clean metal baritone, dark brooding sustained tone, gothic gravity | metal | 3.5/2.8/4.7 | 33% | 4.7/4.2/4.0 | 67% | +33 | solo fail | 4 genres: industrial_shout, noh_utai, vaudeville | ||
| Operatic baritone operatic baritone, dark resonant chest voice, noble sustained line | classical | 4.0/1.5/5.0 | 67% | 3.3/4.3/5.0 | 50% | -17 | solo fail | 4 genres: jazz_scat, rock_classic, tango | ||
| Appalachian ballad unaccompanied appalachian ballad, plain hard modal tone, ornamented free r | folk | 2.8/3.0/5.0 | 17% | 3.5/1.8/5.0 | 50% | +33 | FAIL | 3 genres: blues_classic_female, cantonese_opera | ||
| Roots reggae roots reggae, nasal plaintive tone, off-beat phrasing, prophetic delivery | urban | 4.7/4.5/5.0 | 100% | 4.0/5.0/4.5 | 50% | -50 | solo fail | 3 genres: blues_chicago, salsa_sonero | ||
| Saeta spanish saeta, unaccompanied processional cry from a balcony, raw and pier | world | 3.5/1.0/4.7 | 50% | 3.8/1.2/3.8 | 40% | -10 | FAIL | 3 genres: chanson, torch_song | ||
| Playground chant children's playground chant, sing-song taunting cadence, shouted and rhyth | functional | 4.2/2.5/4.7 | 50% | 4.2/2.3/5.0 | 67% | +17 | FAIL | 3 genres: child_folk, nursery_rhyme_child | ||
| Lovers rock lovers rock, sweet gentle reggae-soul, tender high register | urban | 4.3/4.5/5.0 | 67% | 4.0/3.5/4.2 | 33% | -33 | FAIL | 3 genres: jazz_standard, scots_gaelic_waulking | ||
| German Lied art-song lieder delivery, intimate recital voice, careful diction, restrai | classical | 2.7/2.2/4.8 | 17% | 4.7/2.5/5.0 | 83% | +67 | genre fail | 3 genres: muezzin_style, qawwali | ||
| Black metal shriek black metal shriek, high rasping strangled screech | metal | 2.5/0.8/4.5 | 17% | 4.2/3.3/4.7 | 67% | +50 | FAIL | 3 genres: metal_death, turkish_uzun_hava | ||
| Singer-songwriter singer-songwriter, plainspoken confessional delivery, small honest untrain | folk | 3.5/4.5/4.7 | 50% | 4.3/5.0/5.0 | 83% | +33 | PASS | 2 genres: post_rock_wail | ||
| Patter song comic patter song, extremely fast precise consonants, breathless wit | theatre | 3.0/1.0/4.3 | 33% | 5.0/1.3/4.7 | 83% | +50 | genre fail | 2 genres: auctioneer | ||
| Tuvan throat singing tuvan khoomei throat singing, low fundamental drone with whistling overton | world | 3.0/1.2/4.5 | 17% | 3.5/0.7/3.8 | 33% | +17 | FAIL | 2 genres: bluegrass_high_lonesome | ||
| Delta blues delta blues, rough weathered voice, bent notes, moaning field-holler infle | blues | 3.5/1.8/4.7 | 50% | 4.2/4.5/5.0 | 67% | +17 | solo fail | 2 genres: cowboy_yodel | ||
| Dancehall dancehall, rapid patois toasting sliding into melody, rhythmic and percuss | urban | 3.7/1.3/5.0 | 50% | 3.7/2.2/4.7 | 17% | -33 | FAIL | 2 genres: boy_treble | ||
| Cabaret cabaret, arch knowing delivery, spoken asides sliding into sung line | theatre | 3.3/4.0/5.0 | 50% | 3.7/3.7/5.0 | 50% | 0 | FAIL | 2 genres: maori_waiata | ||
| Ranchera mexican ranchera, huge open-throated grito, proud sustained belt | world | 4.5/1.7/5.0 | 83% | 4.0/1.0/4.2 | 50% | -33 | FAIL | 2 genres: country_classic | ||
| Emo emo, cracking earnest voice, whiny nasal edge tipping into desperate yells | rock | 3.2/3.8/5.0 | 33% | 3.8/5.0/5.0 | 50% | +17 | solo fail | 2 genres: teen_choir_solo | ||
| Hindustani khayal hindustani khayal, slow raga exposition, microtonal meend glides, sustaine | world | 3.0/0.8/5.0 | 33% | 4.5/0.8/4.5 | 67% | +33 | FAIL | 2 genres: ghazal | ||
| J-pop j-pop, high clear cute-inflected tone, wide melodic leaps, earnest | pop | 3.2/4.7/5.0 | 33% | 3.2/4.8/4.5 | 17% | -17 | solo fail | 2 genres: teen_bedroom | ||
| Gothic metal soprano gothic symphonic metal soprano, operatic vibrato over heaviness, ethereal | metal | 3.2/4.5/5.0 | 33% | 4.0/2.2/4.2 | 50% | +17 | FAIL | 2 genres: opera_sop_dram | ||
| Sea shanty sea shanty, hearty rolling call, blunt communal projection | folk | 3.3/3.8/5.0 | 50% | 3.8/4.3/5.0 | 67% | +17 | solo fail | 2 genres: mongolian_urtiin | ||
| Glam rock glam rock, theatrical androgynous sneer, camp vibrato | rock | 4.0/4.5/4.5 | 50% | 2.5/5.0/5.0 | 17% | -33 | solo fail | 2 genres: work_song | ||
| Afrobeats afrobeats, lilting melodic pidgin-inflected phrasing, relaxed buoyant swin | urban | 3.5/3.2/4.7 | 33% | 3.8/4.8/5.0 | 50% | +17 | solo fail | 1 genres: | ||
| Bel canto bel canto, effortless even legato across registers, elegant messa di voce | classical | 3.5/2.7/4.8 | 50% | 2.5/2.7/5.0 | 17% | -33 | FAIL | 1 genres: | ||
| Bolero latin bolero, velvet romantic croon, dramatic sustained romantic line | world | 4.0/2.3/5.0 | 67% | 4.2/3.2/4.5 | 50% | -17 | FAIL | 1 genres: | ||
| Britpop britpop, adenoidal regional-accented drawl, cocky flat delivery | rock | 4.0/3.2/5.0 | 67% | 4.5/4.0/4.5 | 83% | +17 | PASS | 1 genres: | ||
| Chorister solo solo chorister boy, pure vibrato-free treble, echoing cathedral clarity | classical | 4.0/1.8/5.0 | 67% | 4.0/2.8/5.0 | 67% | 0 | FAIL | 1 genres: | ||
| Doo-wop lead doo-wop lead, sweet earnest tenor, swooping falsetto leaps, vintage warmth | soul | 3.8/4.4/5.0 | 60% | 3.7/4.7/5.0 | 50% | -10 | solo fail | 1 genres: | ||
| Pure falsetto pure falsetto, light hooty unsupported head voice, fragile | functional | 2.8/3.7/5.0 | 17% | 4.3/4.5/5.0 | 83% | +67 | PASS | 1 genres: | ||
| Forró forro, warm northeastern brazilian accent, rustic bouncing delivery | world | 3.5/1.2/4.8 | 50% | 5.0/1.7/4.7 | 83% | +33 | genre fail | 1 genres: | ||
| Grunge grunge, strained throaty rasp, apathetic verses erupting into raw shouting | rock | 4.4/3.6/5.0 | 80% | 5.0/4.8/4.0 | 67% | -13 | solo fail | 1 genres: | ||
| Indie pop indie pop, slightly nasal untrained charm, close-mic breathiness, casual | pop | 2.8/4.7/5.0 | 17% | 3.3/4.8/5.0 | 50% | +33 | solo fail | 1 genres: | ||
| Power metal power metal, soaring operatic high tenor, sustained ringing screams | metal | 3.3/3.0/4.2 | 33% | 4.2/5.0/5.0 | 67% | +33 | solo fail | 1 genres: | ||
| Boy-band tenor boy-band lead, light high tenor, earnest and smooth, adolescent brightness | pop | 1.8/3.7/5.0 | 0% | 4.5/5.0/4.8 | 83% | +83 | PASS | 1 genres: | ||
| Quiet storm quiet storm R&B, smooth low seductive croon, silky sustained falsetto | soul | 3.3/4.8/4.8 | 33% | 3.8/4.8/4.7 | 50% | +17 | solo fail | 1 genres: |
Tango.code2sound's
decode_audio(...)[0], which silently returns one song for a whole batch (the vocoder
therefore runs at batch 1 here).torch.Generator, so a
row's audio is a function of (prompt, sampling params, seed) alone and not of its batch-mates or
its position in the batch. Verified bit-identical for the same genome at two positions. Without
this, two candidates in the same batch are not comparable.