Vocal Music Whisper — ten held-out clips

Each clip was held out of training entirely. Play it, read what each checkpoint says, then compare against the reference caption at the bottom of the card.

✅ The default is v2/unbalanced-lr3e5 — and it is the repo root

Download it → huggingface.co/TTS-AGI/vocal-music-whisper  —  from_pretrained("TTS-AGI/vocal-music-whisper"), no subfolder=. Inference code, training code and the raw judgements are in the same repo.

Chosen on caption quality: Gemini 3.6 Flash was given the audio and the caption and scored the caption on detail, factual correctness and overall impression, over all 158 held-out clips, twice. Overall impression 3.47 / 5 against 3.29 for the balanced retrain, 3.23 for v1 and 2.48 for the base model.

Its lead over the balanced retrain is inside the noise (+0.18, 95 % CI [-0.04, +0.39], p = 0.09). The tie is broken by the axes that do separate — validation loss, voice gender and flagged factual errors — on none of which is it behind. Details in the two tables below.

Whisper fine-tuned to describe singing rather than transcribe it: voice count, register, technique, timbre, emotion.

The base model is laion/music-whisper — LAION's music captioner, itself a whisper-small topology. An earlier version of this page compared against openai/whisper-small; that was wrong, whisper-small is not what these checkpoints were fine-tuned from, and it has been removed along with the claim about lyric transcription that was measured on it.

What laion/music-whisper actually does, measured on all 158 held-out clips: 158 of 158 outputs (100%) are music description rather than lyrics, which is what it was built for — but it describes them as songs, reaching for instrumentation, tempo and production on material that is unaccompanied solo voice. 0 of 158 are lyric transcripts and 0 are degenerate (empty or one phrase repeated to the token cap). Base output averages 137 words. The fine-tune is not teaching it to stop transcribing; it is teaching it to describe the voice instead of an imagined band.

checkpointwhat changedval loss
basenone — laion/music-whisper3.2411
v1 frozen encoder (published)3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices male0.9461
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b>10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo root0.9597
v2 balanced, lr 1e-410 epochs, 50/50 gender-balanced corpus — the retrain0.9714
bedroom pop proven voice: female solo yes · backing no · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music that commences with a gentle acoustic guitar melody. The tempo is slow, approximately 70 beats per minute. A female vocalist then enters, delivering the vocal line in a clear and slightly breathy style. The vocal profile is that of an adult female, characterized by a clear, gentle, and slightly breathy timbre. There are no discernible non-lyrical vocal sounds. The genre of the music is categorized as either Acoustic Pop or Singer-Songwriter. The overall mood of the piece is gentle, reflective, and carries a slight melancholic undertone. The combination of the acoustic guitar and the clear vocal delivery creates a sense of intimacy and vulnerability within the music. This track would be well-suited for a quiet, introspective scene in a film, a coffee shop playlist, or as background music for a coffee shop setting.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo female voice performing an intimate indie-pop melody with a breathy, conversational delivery. The singer employs a light mixed voice with a gentle, narrow vibrato and smooth legato phrasing, occasionally using subtle glottal attacks for emphasis. Her timbre is warm and resonant yet possesses a fragile, airy quality that enhances the melancholic and nostalgic mood. The vocal range spans a comfortable mid-register, characteristic of a mezzo-soprano, covering a modest range that feels both grounded and expressive. The performance is entirely unaccompanied, with no audible instrumentation or background noise, captured in a dry, close-miked setting that emphasizes the singer's proximity and emotional nuance.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo female voice performing an intimate indie-pop melody without any instrumental accompaniment. The singer employs a breathy, delicate delivery with a light mixed voice that occasionally breaks into a delicate falsetto for higher melodic peaks. Her timbre is warm and resonant yet possesses a fragile, airy quality that enhances the song's melancholic and nostalgic mood. The vocal range is primarily in the mezzo-soprano register, moving fluidly between a grounded chest voice and a light, airy head voice. The phrasing is predominantly legato, with subtle glottal onsets and a narrow, controlled vibrato that adds a sense of vulnerability to the performance. The recording is dry and close-miked, capturing the fine details of the vocal texture, with a light touch of reverb that provides a sense of space without masking the detail. The overall expressive arc is one of quiet reflection and longing, maintaining a consistent, intimate character throughout the clip.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single solo female voice performing in an intimate indie-pop style, characterized by a breathy and delicate delivery. The singer employs a light mixed voice with a gentle, narrow vibrato and smooth legato phrasing, occasionally using soft glottal attacks for emphasis. Her timbre is warm and resonant yet possesses a fragile, airy quality that perfectly conveys a sense of melancholy and quiet longing. The vocal range is centered in the mid-to-upper register, suggesting a mezzo-soprano voice type. The performance is entirely unaccompanied, with no instrumental backing or background noise audible. The recording is dry and close-miked, capturing the subtle nuances of the singer's breath and articulation, with a hint of natural room ambience that enhances the sense of proximity.
reference caption (gemini 3 flash, the training target)This recording features a single solo female voice, unaccompanied and captured in a very dry, close-mic'd studio environment. The performance is rooted in an indie-pop aesthetic, characterized by a breathy, intimate vocal style that utilizes a light mixed voice and transitions into a pure, fragile head voice on the ascending melodic lines. The singer's timbre is warm and resonant yet maintains an airy quality throughout, with a subtle, controlled vibrato and smooth legato phrasing. There are no backing vocals, harmonies, or instrumental accompaniments present, allowing the nuanced glottal attacks and delicate emotional arc—ranging from melancholic longing to gentle resignation—to remain the focal point. The recording is clean and polished, with no audible background noise or digital artifacts.
brill building pop sixties voice: female solo yes · backing yes · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music that immediately establishes a relaxed and introspective mood. The primary vocal element is a female vocalist, whose voice is characterized by a clear and slightly breathy timbre. The vocal performance is devoid of any non-lyrical sounds. The song's structure is built upon a simple, repetitive melody played on a synthesizer. The tempo is slow, contributing to the overall feeling of ease. The genre is categorized as either Indie Pop or Dream Pop. The mood is primarily introspective, with a subtle undercurrent of melancholy. The combination of the simple melody and the breathy vocal delivery creates a sense of intimacy and vulnerability within the music. This track would be well-suited for a quiet, reflective scene in a film, a late-night listening session, or as background music for a vlog.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male voice that begins with an intimate, breathy delivery characteristic of contemporary indie folk. As the track progresses, the arrangement expands to include lush, multi-tracked backing harmonies that create a rich, choral texture. The singer employs a smooth legato phrasing with subtle glottal attacks and a controlled vibrato that widens during sustained notes. The vocal timbre is warm and resonant, transitioning from a fragile, airy quality in the lower register to a more powerful, slightly gritty tone as the intensity builds. The performance spans a comfortable baritone-tenor range, moving from a grounded chest voice to a soaring, emotive upper register. The expressive arc moves from a sense of quiet determination to a defiant, anthemic climax. This is an entirely unaccompanied vocal performance with no audible instrumentation or background noise. The recording is polished and professional, featuring a close-mic'd proximity that captures the fine details of the vocal texture, enhanced by a tasteful amount of studio reverb that adds depth without obscuring the clarity.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male voice that begins with an intimate, breathy delivery characteristic of contemporary indie folk. The singer employs a warm, resonant timbre with a slight huskiness, utilizing a mix of chest and head voice to navigate a mid-to-high baritone range. As the clip progresses, the solo line is joined by layered backing vocals and lush harmonies, creating a rich, choral texture that enhances the emotional intensity. The vocal technique shifts from a delicate, almost whispered delivery to a powerful, belted mixed voice in the higher register, where the singer employs subtle glottal attacks and a controlled vibrato to convey a sense of yearning and determination. The recording is entirely unaccompanied, with no audible instrumentation or background noise, presenting a clean and focused vocal performance. The production is polished and dry, with a close-miked quality that captures the fine details of the singer's breath and articulation, enhanced by a subtle, transparent reverb that adds a sense of space without obscuring the clarity of the performance.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a powerful male solo voice that begins in a resonant baritone register before expanding into a rich, multi-layered arrangement with backing harmonies. The performance is rooted in the indie-folk genre, characterized by a soulful and defiant emotional arc. The singer utilizes a strong mixed voice and resonant belting, particularly in the higher passages, where a controlled vibrato and subtle melismatic flourishes add to the expressive delivery. The timbre is warm and slightly gritty, conveying a sense of resilience and determination. While the focus is on the vocals, there is a clear rhythmic accompaniment of finger snaps and a subtle, atmospheric synth pad providing harmonic depth. The production is polished and professional, with a noticeable studio reverb that enhances the vocal's presence and creates a sense of space.
reference caption (gemini 3 flash, the training target)This a cappella performance begins with a single, resonant male solo voice before expanding into a rich texture of layered backing harmonies and doubled vocal parts. The track sits firmly in the Indie Folk genre, leaning into an Americana aesthetic with its soulful, storytelling delivery. The vocalist demonstrates a versatile technique, transitioning from a warm, conversational chest voice to powerful, high-register belting in a mixed voice. The timbre is earthy and resonant, with a subtle grit that emerges during more intense passages, complemented by a controlled vibrato and expressive glottal attacks. The range spans from a comfortable baritone mid-range to a soaring tenor belt, conveying a defiant and resilient emotional arc that builds in intensity. The recording is entirely unaccompanied, featuring a dry, close-miked character with a tasteful amount of studio reverb that adds depth without obscuring the vocal detail. The performance is polished and clean, with no audible background noise or instrumentation, allowing the pure vocal textures and intricate ornaments to take center stage.
melodie francaise classical voice: female solo yes · backing yes · instruments yes
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music primarily characterized by a female vocalist. The vocal performance is delivered by an adult female, exhibiting a clear and slightly nasal timbre. The vocal quality is described as clear and slightly nasal. The instrumentation is minimalistic, consisting solely of a simple acoustic guitar melody. The tempo of the music is slow, estimated to be around 70 beats per minute. The genre of the music is categorized as either Indie Folk or Singer-Songwriter. The overall mood of the piece is melancholic and reflective. The simple instrumentation, combined with the clear vocal delivery, contributes to a sense of intimacy and vulnerability within the music. This track would be well-suited for a quiet, introspective scene in a film, or as background music for a coffee shop setting.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male tenor voice that begins with an intimate, breathy delivery before expanding into a rich, multi-tracked arrangement with layered harmonies and doubled vocal lines. The genre is contemporary pop, characterized by a soulful and emotive vocal style. The singer employs a mix of chest and head voice, transitioning into a powerful belt during the more intense passages. Notable techniques include subtle melismatic flourishes, a controlled vibrato, and occasional glottal attacks that add a sense of urgency. The timbre is warm and resonant, though it takes on a slightly gritty, strained quality during the high-energy peaks to convey deep yearning and nostalgia. The recording is entirely unaccompanied, with no instrumental backing or background noise, and it possesses a polished, studio-quality sound with a moderate amount of reverb that adds depth and space. The expressive arc moves from a quiet, reflective opening to a passionate, high-register climax, showcasing a wide dynamic range and a seamless transition between registers.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male voice that begins with a conversational, intimate delivery before being joined by layered backing vocals and lush harmonies that create a powerful, anthemic sound. The genre is firmly rooted in Indie Folk, characterized by a warm, resonant timbre and a slight grit that adds emotional weight to the performance. The singer demonstrates a versatile range, moving from a breathy lower register to a soaring, belted mixed voice in the higher passages. The vocal technique includes subtle glottal attacks and a controlled vibrato that widens during sustained notes. The expressive arc is one of building intensity, starting with a quiet, nostalgic reflection and culminating in a passionate, yearning climax. The recording is entirely unaccompanied, featuring only the human voice with no instrumental backing or background noise. The production is polished and dry, with a close-miked quality that is enhanced by a touch of studio reverb, providing a sense of space without masking the detail of the vocal performance.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a solo male tenor voice that begins with an intimate, breathy delivery before building into a powerful, multi-layered vocal arrangement. The genre is firmly rooted in contemporary pop-rock, characterized by a transition from a warm, resonant chest-mix in the lower register to a soaring, belted upper register. The singer employs a controlled vibrato and smooth legato phrasing, with occasional glottal attacks that add emotional weight to the performance. As the clip progresses, the vocal texture thickens significantly as lush backing harmonies and layered parts are introduced, creating a rich, choral effect. The expressive arc moves from a quiet, nostalgic reflection to a passionate, yearning climax. The recording is entirely unaccompanied, with no instrumental backing or background noise. It has a polished, studio-quality feel with a noticeable amount of reverb that adds depth and space, particularly during the multi-tracked sections.
reference caption (gemini 3 flash, the training target)This recording features a solo male tenor performing an acoustic pop ballad, accompanied by a crisp, rhythmic acoustic guitar. The singer employs a versatile technique, transitioning from a warm, conversational chest voice in the lower register to powerful, resonant belting in the upper-mid range. The expressive arc begins with a nostalgic, breathy quality and builds into a passionate, high-energy climax characterized by wide vibrato and intense glottal attacks. As the track reaches its peak, layered harmony parts emerge, creating a rich, choral texture that enhances the emotional weight of the performance. The timbre is generally clear and resonant, though it takes on a slight, expressive grit during the more forceful belted passages. The recording is polished and well-produced, with a subtle room reverb that adds depth, concluding with a series of intricate melismatic runs that showcase the singer's agility.
lantern s1234 rep other voice: female solo yes · backing no · instruments yes
base — laion/music-whisper, no fine-tuningThe listener hears a track that immediately establishes a foundation with a simple, repeating piano melody. The tempo is slow, approximately 70 beats per minute. A female vocalist then enters, delivering the vocal line with a breathy and intimate style. The vocal profile is that of an adult female, characterized by a soft, breathy, and intimate timbre. The vocalist also incorporates non-lyrical sounds, specifically soft sighs and breaths, which contribute to the overall atmosphere. The genre is classified as either Lo-fi or Indie Pop. The mood of the piece is predominantly melancholic, introspective, and intimate. The combination of the simple piano melody and the breathy vocals creates a sense of vulnerability and emotional depth. The track's characteristics suggest it would be well-suited for a quiet, reflective scene in a film, a late-night listening session, or as background music for a sad or emotionally charged moment.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male voice, likely a tenor, performing an intimate indie-pop melody with a warm and slightly breathy timbre. The singer employs a mix of chest and head voice, transitioning into a delicate falsetto for the higher melodic leaps. His technique is characterized by smooth legato phrasing and a narrow, controlled vibrato that adds a sense of vulnerability to the performance. The vocal range is centered in the mid-register, conveying a reflective and melancholic emotional arc that shifts toward a more hopeful and fragile tone. While the focus is on the voice, a rhythmic finger-snap or percussive click provides a steady tempo throughout the clip. The recording is dry and close-miked, capturing the subtle nuances of the singer's breath and articulation, with a light touch of reverb that adds a sense of space without obscuring the clarity of the performance.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo female voice that begins with an intimate, breathy delivery, characteristic of contemporary indie-pop. The singer employs a light mixed voice with a gentle, controlled vibrato and smooth legato phrasing, occasionally using subtle glottal attacks for emotional emphasis. Her timbre is warm and resonant, sitting comfortably in a mezzo-soprano range, with a clear, pure tone that occasionally leans into a delicate breathiness for expressive effect. At the midpoint, the texture thickens as layered backing vocals and lush harmonies enter, creating a rich, choral-like arrangement that supports the lead's soaring melodic lines. The performance conveys a sense of nostalgic longing and quiet resolution, with an expressive arc that builds in intensity towards the end of the clip. The recording is entirely unaccompanied, featuring only the voice with no instrumental backing or background noise. The production is polished and dry, with a close-miked quality and a subtle touch of reverb that adds depth without obscuring the clarity of the vocal layers.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a solo female voice performing in a contemporary indie-pop style, characterized by a warm and intimate timbre. The singer employs a breathy, delicate delivery, utilizing a mix of head voice and light chest resonance to navigate a mid-to-high register typical of a mezzo-soprano. Her phrasing is predominantly legato, with subtle glottal attacks and a gentle, narrow vibrato that adds to the emotional vulnerability of the performance. The expressive arc is one of quiet reflection and longing, conveyed through a fragile yet controlled vocal texture. While the voice is the primary focus, a rhythmic finger-snap or light percussive click provides a steady tempo throughout the clip. The recording is dry and close-miked, capturing the fine details of the singer's breath and articulation, with a hint of natural room ambience and minimal processing.
reference caption (gemini 3 flash, the training target)A single solo female voice, possessing a mezzo-soprano range, delivers a performance rooted in the indie-pop and contemporary folk genres. The vocal technique is marked by a breathy, intimate approach, utilizing a warm and resonant timbre that occasionally thins into a fragile, pure head voice. Her phrasing is predominantly legato, featuring gentle glottal entries and a subtle, narrow vibrato that enhances the reflective and tender emotional arc. The recording is not entirely unaccompanied; a muffled, low-frequency rhythmic pulse is audible in the background, providing a minimalist percussive foundation. The production is characterized by a significant amount of hall-like reverb, creating a wide and atmospheric stereo field, while the vocal remains well-compressed and clear, suggesting a polished studio environment.
metal symphonic soprano metal voice: female solo yes · backing yes · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music primarily characterized by a female vocalist. The vocal performance is delivered by an adult female, exhibiting a clear and slightly nasal timbre. The vocal quality is notably clear, with a subtle vibrato adding depth to the sound. The instrumentation is minimal, consisting solely of a simple acoustic guitar melody. The tempo of the music is slow, estimated to be around 60 beats per minute. The genre of the music is classified as either Folk or Singer-Songwriter. The overall mood conveyed by the music is one of gentleness, nostalgia, and a touch of melancholy. The combination of the simple acoustic guitar melody and the clear vocal delivery creates a sense of intimacy and vulnerability within the listener. This musical piece would be well-suited for a quiet, reflective scene in a film, or as background music for a coffee shop, or even for a relaxing evening at home.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male voice, likely a baritone or high tenor, performing a folk-style lullaby with a gentle, storytelling quality. The singer employs a smooth legato phrasing, transitioning seamlessly between a resonant chest-dominant mix and a lighter, more breathy head voice in the upper register. A subtle, narrow vibrato is used sparingly at the ends of phrases, adding a touch of warmth to the otherwise intimate and pure tone. The vocal delivery is characterized by clear diction and a light, expressive placement that conveys a sense of calm and comfort. The recording is entirely unaccompanied, with no instrumental backing or background noise, and the sound is dry and close-miked, providing a high level of detail and presence.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male voice that begins with a gentle, narrative quality, eventually joined by layered backing vocals and lush harmonies that create a rich, choral texture. The performance sits firmly within the contemporary folk or musical theater genre, characterized by its storytelling approach and clear, articulate diction. The singer employs a smooth legato phrasing with a controlled, medium-width vibrato that adds warmth to the sustained notes. The vocal timbre is resonant and warm, possessing a pure, pure quality that remains consistent across a baritone-tenor range. The expressive arc moves from a soothing, lullaby-like opening to a more rhythmic and harmonically dense conclusion. This is an entirely unaccompanied vocal performance with no audible instrumentation or background noise. The recording is polished and dry, with a subtle room reverb that provides a sense of space while maintaining a close, intimate feel.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single solo male voice performing a folk-inspired lullaby with a gentle, narrative quality. The singer employs a smooth legato phrasing, characterized by a warm and resonant timbre that remains consistent throughout the performance. His technique is marked by a controlled, medium-width vibrato and a clear, pure tone that avoids excessive ornamentation. The vocal register sits comfortably in the baritone-tenor range, moving fluidly through a series of melodic leaps. The expressive arc is one of soothing comfort and nostalgia, conveyed through a steady, gentle delivery. The recording is entirely unaccompanied, with no instrumental backing or background noise. The audio is dry and close-miked, providing an intimate and polished sound with very little room ambience or artificial reverb.
reference caption (gemini 3 flash, the training target)This a cappella performance features a solo male voice, likely a high baritone or tenor, singing a piece that blends contemporary folk with a musical theatre sensibility. The singer employs a warm, resonant timbre with a smooth legato delivery, transitioning from a gentle, breathy head-mix in the lower register to a powerful, belted mixed voice as the melody ascends. A consistent, controlled vibrato adds depth to sustained notes, while subtle glottal attacks and occasional ornaments provide a sense of folk-inspired storytelling. Midway through the clip, the texture thickens as additional vocal layers are introduced, including a soaring, high-register counter-melody that creates a rich, choral effect. The expressive arc begins with a soothing, lullaby-like quality and builds into a triumphant, emotive climax, conveying a sense of warmth and protection. The recording is entirely unaccompanied, with no instrumental backing or background noise, and is captured with a dry, close-mic intimacy enhanced by a subtle, polished studio reverb.
americana singer songwriter proven voice: male solo no · backing yes · instruments yes
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music primarily characterized by a male vocalist. The vocal performance is delivered by an adult male, exhibiting a clear and slightly nasal timbre. The vocal quality is notably clear, with a subtle nasal resonance. The instrumentation is minimal, consisting solely of a simple acoustic guitar melody. The tempo of the music is slow, estimated to be approximately 60 beats per minute. The overall sonic texture is sparse, contributing to a sense of intimacy. The genre of the music is classified as either Folk or Singer-Songwriter. The mood evoked by the music is melancholic and reflective. The simplicity of the acoustic guitar melody, combined with the clear vocal delivery, fosters a sense of intimacy and vulnerability within the listener. This musical piece would be well-suited for a quiet, introspective scene in a film, or as background music for a coffee shop setting.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male baritone performing a folk-style ballad with a theatrical, storytelling quality. The singer employs a resonant mixed voice, transitioning into a powerful belt during the more intense passages. His timbre is warm and rich, characterized by a consistent, medium-width vibrato and clear, precise diction. The phrasing is predominantly legato, though he uses subtle glottal attacks for emphasis on certain words, adding a sense of urgency and emotional weight to the performance. The vocal range spans approximately an octave and a half, showcasing a well-supported chest-dominant mix. The expressive arc begins with a reflective, almost somber tone and builds toward a more defiant and passionate climax. The recording is entirely unaccompanied, with no instrumental backing or background noise audible. The audio is dry and close-miked, with a touch of artificial reverb that adds depth without obscuring the clarity of the vocal delivery.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male baritone performing a dramatic musical theater piece with a powerful, resonant delivery. The singer employs a strong mixed voice and belting technique, particularly in the higher register where the tone becomes more intense and slightly gritty. His phrasing is primarily legato, punctuated by clear glottal attacks and a controlled, medium-width vibrato that adds emotional weight to the sustained notes. The timbre is warm and full-bodied, conveying a sense of yearning and determination that builds throughout the clip. The recording is entirely unaccompanied, with no instrumental backing or background noise, and the sound is dry and close-miked with a subtle touch of reverb that enhances the natural resonance of the voice.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a solo male tenor voice that is eventually joined by layered backing vocals and harmonies, creating a rich, choral texture. The genre is a blend of contemporary musical theater and pop-rock, characterized by a powerful, resonant vocal delivery. The singer employs a strong mixed voice and transitions into a belt, particularly on the higher sustained notes. His technique includes a controlled vibrato and smooth legato phrasing, with occasional glottal attacks for emphasis. The timbre is warm and clear, with a bright resonance that remains consistent across a range of approximately an octave and a half. The expressive arc begins with a sense of quiet determination and builds into a triumphant, anthemic climax. The recording is entirely unaccompanied, with no instrumental backing or background noise. The production is polished and dry, with a close-miked intimacy and a subtle touch of reverb that enhances the vocal's presence.
reference caption (gemini 3 flash, the training target)This recording features a male lead vocalist accompanied by an acoustic guitar, with prominent multi-part vocal harmonies appearing during the chorus sections. The genre is contemporary indie folk, characterized by an earnest and storytelling-driven delivery. The lead singer, likely a baritone or high baritone, utilizes a well-supported mixed voice with a warm, resonant timbre that occasionally takes on a slightly gritty edge for emotional emphasis. His technique includes subtle vibrato at phrase endings and occasional glottal attacks that add to the intimate, conversational feel of the verses. The vocal range is moderate, staying mostly within a comfortable mid-register. The backing vocals provide rich, layered harmonies that expand the sonic texture, creating a sense of communal strength against the more solitary verses. The overall emotional arc moves from a quiet, reflective melancholy to a more determined and resilient tone. The recording is polished and clear, with a dry, close-mic'd quality on the lead vocal and a light, natural-sounding reverb that gives the performance a sense of space. The acoustic guitar is clearly audible throughout, providing a rhythmic and harmonic foundation.
metal doom clean metal voice: male solo no · backing yes · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a recording dominated by a male vocalist. The vocal performance is characterized by a mature adult male voice, exhibiting a clear and slightly nasal timbre. The vocal delivery is notably expressive, conveying a sense of emotion. The recording contains no non-lyrical vocal sounds. The vocal style is consistent with a folk or country music genre. The overall mood of the piece is melancholic and reflective. The instrumentation is simple, featuring a solo male vocalist. The recording's production quality is relatively clean, though it lacks significant polish. This piece would be suitable for a scene in a film or television show depicting a character's internal struggles.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a powerful male tenor performing in a contemporary musical theatre style, beginning with a solo vocal that soon expands into a rich, multi-layered arrangement with backing harmonies. The singer employs a robust mixed-voice technique, transitioning into a resonant belt with a controlled, medium-width vibrato and precise legato phrasing. His timbre is warm and resonant, possessing a bright, forward placement that ensures clarity across a wide range that showcases both a solid chest-dominant mix and a soaring upper register. The expressive arc is one of growing determination and triumph, conveyed through a passionate delivery that remains consistent throughout the clip. The audio is entirely unaccompanied, featuring only the voice with no instrumental backing or background noise. The production is polished and professional, with a noticeable hall-like reverb that adds depth and a sense of space, while the stereo field widens significantly as the harmony parts enter, creating a wide and immersive soundstage.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male baritone performing in a contemporary musical theatre style, characterized by a resonant and warm timbre. The singer employs a powerful mixed voice with a controlled vibrato and clear, crisp diction, moving from a steady legato in the verses to a more rhythmic, declamatory delivery in the chorus. His range spans from a grounded chest register to a bright, ringing upper-mid range, showcasing a high baritone or dramatic tenor voice type with a bright, forward placement. The performance conveys a sense of determination and emotional intensity, with an expressive arc that builds from a focused, narrative opening to a more expansive and passionate climax. The audio is entirely unaccompanied, with no instrumental backing or background noise, and the recording is dry and close-miked, providing an intimate and polished sound with a subtle touch of reverb to add space.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single solo male voice performing in a soulful, gospel-influenced style, delivered with a high baritone or tenor range. The singer employs a powerful mixed voice, transitioning into a resonant belt for the more emotive passages. His technique is characterized by a controlled, medium-width vibrato and occasional glottal attacks that add a sense of urgency and grit to the performance. The timbre is warm and rich, yet it carries a certain raw, soulful edge that conveys a deep sense of hope and determination. The vocal range is moderate, focusing on the mid-to-upper register, with the expressive arc building in intensity towards the end of the clip. The recording is entirely unaccompanied, with no instrumental backing or background noise. The audio quality is polished and dry, with a close-miked perspective that captures the intimate details of the vocal performance, though a subtle room reverb is audible to provide a sense of space.
reference caption (gemini 3 flash, the training target)This soulful a cappella performance features a lead male tenor whose voice is eventually joined by layered backing harmonies and rhythmic vocalizations. The singer employs a powerful belting technique in the upper register, characterized by a controlled vibrato and precise glottal attacks that emphasize the lyrical delivery. His timbre is warm and resonant, possessing a slight grit during the more intense passages while maintaining a pure, focused tone in the lower melodic lines. Spanning a wide range from the mid-chest register to a soaring high tenor belt, the performance showcases impressive vocal agility and soulful ornamentation. The expressive arc builds from a contemplative, intimate opening to a triumphant, multi-voiced climax, conveying a sense of deep conviction and emotional strength. The recording is entirely unaccompanied, with no instrumental backing or background noise, and features a polished studio quality with a subtle, natural reverb that enhances the vocal depth.
choir other voice: male solo yes · backing no · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music that commences with a gentle acoustic guitar melody. The tempo is slow, approximately 70 beats per minute. A male vocalist then enters, delivering the vocal part. The vocal profile features an adult male voice, characterized by a clear, slightly raspy, and sincere timbre. There are no non-lyrical vocal sounds present. The genre of the music is categorized as either Folk or Singer-Songwriter. The overall mood of the piece is reflective, hopeful, and tinged with a slight melancholic undertone. The acoustic guitar's presence, coupled with the sincere vocal delivery, contributes to a sense of intimacy and vulnerability within the music. This track would be well-suited for a quiet, introspective scene in a film, or as background music for a coffee shop setting.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleA single male solo voice performs a contemporary folk melody with a warm, resonant timbre and a gentle, breathy quality. The singer employs a smooth legato phrasing, utilizing a mixed voice that transitions seamlessly between a grounded chest register and a light, airy head voice. Subtle vibrato is used sparingly at the ends of phrases, while the overall delivery is intimate and reflective. The performance is entirely unaccompanied, with no instruments or backing vocals present. The recording is dry and close-miked, capturing the natural resonance of the voice with a hint of natural room ambience, resulting in a polished and clear sound.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootA solo male baritone performs an intimate, a cappella folk melody with a warm and resonant timbre. The vocal delivery is characterized by a breathy, mixed-voice technique and a gentle, controlled vibrato that adds a sense of vulnerability. The phrasing is predominantly legato, with subtle glottal attacks and a slight breathiness that enhances the emotional intimacy of the performance. The recording is dry and close-miked, capturing the fine details of the singer's breath and articulation without any instrumental accompaniment or digital effects. The overall mood is one of quiet reflection and longing, maintained through a consistent expressive arc.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single male solo voice performing a folk-inspired melody with a soulful, intimate delivery. The singer employs a smooth legato phrasing, characterized by gentle glottal attacks and a subtle, narrow vibrato that adds a sense of vulnerability to the performance. His voice is a warm, resonant baritone, moving fluidly through a mid-range register that feels comfortable and effortless. The tone is primarily pure and clear, though it carries a slight breathiness that enhances the emotional intimacy of the piece. The expressive arc is one of quiet reflection and longing, conveyed through a steady, controlled delivery. The recording is entirely unaccompanied, with no instrumental backing or background noise. It has a dry, close-miked quality with a hint of natural room reverb, providing a polished and intimate listening experience.
reference caption (gemini 3 flash, the training target)A single male solo voice performs an intimate folk-style melody with a gentle, storytelling delivery. The singer employs a breathy, light-toned mixed voice featuring a subtle, narrow vibrato at the end of sustained phrases. The phrasing is predominantly legato, with soft glottal attacks that enhance the emotional vulnerability of the performance. The timbre is warm and resonant, yet possesses a delicate, fragile quality throughout the mid-to-high register, suggesting a light baritone or lyric tenor voice type. The recording conveys a sense of quiet longing and serenity, maintaining a steady, hushed expressive arc. This is an entirely unaccompanied vocal performance with no instrumental backing or audible background noise. The recording character is dry and close-miked, capturing a slight hint of natural room ambience for a polished, professional sound.
brill building pop sixties voice: female solo yes · backing no · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a children's song characterized by its simplicity and innocent nature. The primary vocal element is a young female child, whose voice is clear and slightly nasal. The vocal timbre is described as sweet and innocent. There are no non-lyrical vocal sounds present. The song's lyrics, which are not quoted, are delivered in a way that suggests a focus on the child's voice. The song's genre is classified as Children's Music, specifically designed for a young audience. The overall mood is cheerful, innocent, and playful. The simple melody and the childlike vocal delivery are designed to be easily accessible and easily memorable. The song's structure and style make it suitable for various applications, including children's television shows, preschool classrooms, or as background music for children's activities.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a solo male voice, likely a baritone, performing a folk-style lullaby with a gentle, lullaby-like quality. The singer employs a smooth legato phrasing, characterized by a warm, resonant timbre and a light, controlled vibrato that adds a sense of comfort. Navigating a comfortable mid-range, the vocalist uses a mix of chest and head voice, occasionally transitioning into a delicate falsetto for higher melodic peaks. The performance is entirely unaccompanied, with no instrumental backing or background noise, captured in a dry, close-miked setting that emphasizes the intimate, storytelling nature of the piece. The expressive arc is consistently soothing and calm, maintaining a steady, soothing tone throughout the clip.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo female voice that begins with a gentle, intimate delivery, eventually joined by layered backing vocals and lush harmony parts that create a rich, choral texture. The genre is firmly rooted in contemporary folk, characterized by a pure and resonant vocal tone. The singer employs a smooth legato phrasing with a light, natural vibrato and subtle glottal attacks that add a sense of storytelling. The timbre is warm and pure, with a slight breathiness in the lower register that transitions into a clear, bright head voice as the melody ascends. The range spans a comfortable mid-tenor register, showcasing a well-supported and consistent vocal register. Emotionally, the performance conveys a sense of serene nostalgia and gentle comfort, with an expressive arc that builds in intensity towards the end of the clip. The recording is entirely unaccompanied, featuring only the voice with no instrumental backing or background noise. The production is polished and dry, with a close-miked intimacy and a subtle touch of reverb that provides a sense of space while maintaining a clear, professional sound quality.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single solo female voice performing a folk-style lullaby with a gentle, rhythmic delivery. The singer employs a light, breathy head voice with a narrow vibrato and smooth legato phrasing, occasionally using soft glottal attacks for emphasis. Her timbre is warm and resonant yet possesses a delicate, airy quality that suggests a mezzo-soprano voice type. The performance is entirely unaccompanied, with no instrumental backing or background noise audible. The recording is dry and close-miked, capturing the intimate nuances of the vocal performance with minimal room reverb and a clear, polished character.
reference caption (gemini 3 flash, the training target)A single, resonant male voice performs this folk-inspired lullaby with a tender and soothing delivery. The singer, likely a baritenor, employs a smooth mixed-voice technique, characterized by gentle legato phrasing and a moderate, controlled vibrato that blossoms at the end of sustained notes. The timbre is warm and clear, with a slight breathiness in the lower register that enhances the intimate, storytelling quality of the performance. The vocal range spans roughly an octave and a half, moving effortlessly between a grounded chest register and a light, airy head voice. The expressive arc is subtle, building a gentle intensity in the middle section before subsiding into a quiet, lulling conclusion. This is an entirely unaccompanied vocal recording, captured with a dry, close-miked quality and a hint of natural room reverb that provides a sense of space without any audible instrumentation or background noise.
gregorian chant classical voice: male solo yes · backing no · instruments no
base — laion/music-whisper, no fine-tuningThe listener hears a piece of music that commences with a gentle acoustic guitar melody. The tempo is slow, approximately 70 beats per minute. A male vocalist then enters, delivering the vocal part. The vocal profile features an adult male voice, characterized by a clear, slightly raspy, and sincere timbre. There are no non-lyrical vocal sounds present. The genre of the music is categorized as either Folk or Singer-Songwriter. The overall mood of the piece is reflective, hopeful, and carries a slight melancholic undertone. The acoustic guitar's presence, coupled with the sincere vocal delivery, cultivates a sense of intimacy and vulnerability within the listener. This musical composition would be well-suited for a quiet, introspective scene in a film, a coffee shop playlist, or as background music for a documentary.
v1 frozen encoder (published) — 3 epochs, unbalanced corpus — calls 31 of 34 held-out female voices maleThis recording features a single male solo voice performing in an alternative rock style, characterized by a powerful and emotive delivery. The singer employs a resonant mixed voice with a warm, slightly gritty timbre that adds a sense of raw urgency to the performance. Spanning a mid-to-high baritone range, the vocalist demonstrates excellent control over their mixed register, transitioning seamlessly into a soaring belt during the more intense passages. The phrasing is predominantly legato, punctuated by occasional glottal attacks and subtle melismatic flourishes that enhance the expressive arc, which moves from a grounded, contemplative mood to a defiant and passionate climax. The audio is entirely unaccompanied, with no instrumental backing or background noise, captured in a dry, close-miked setting that emphasizes the intimacy and detail of the vocal performance.
v2 unbalanced, lr 3e-5 &mdash; <b>THE DEFAULT</b> — 10 epochs, original corpus — best caption quality and best voice-gender of any fine-tune; shipped at the repo rootThis recording features a solo male voice that begins with an intimate, breathy delivery in the lower baritone register, characteristic of contemporary indie folk. The singer employs a warm, resonant timbre with a slight huskiness, utilizing a mix of chest and head voice to navigate a mid-range baritone range. As the clip progresses, the texture thickens significantly with the addition of layered backing vocals and lush harmonies, creating a powerful, choral effect that enhances the emotional intensity. The vocal technique is sophisticated, featuring smooth legato phrasing, subtle glottal attacks for emphasis, and a controlled vibrato that widens during the more climactic passages. The expressive arc moves from a quiet, introspective melancholy to a defiant and anthemic resolution. This is an entirely unaccompanied vocal performance with no audible instrumentation or background noise. The recording is polished and dry, with a close-miked proximity that captures the fine details of the vocal texture, enhanced by a subtle, transparent reverb that adds a sense of space without obscuring the clarity of the voices.
v2 balanced, lr 1e-4 — 10 epochs, 50/50 gender-balanced corpus — the retrainThis recording features a single male solo voice performing in a soulful, gospel-influenced style. The singer demonstrates a command of his baritone-tenor range, moving from a warm, resonant chest voice in the lower register to a powerful, belted mixed voice as the melody ascends. His technique is characterized by a controlled vibrato and smooth legato phrasing, punctuated by occasional glottal attacks that add emotional weight to the delivery. The timbre is rich and earthy, with a slight grit appearing as the intensity increases, conveying a sense of defiance and hope. The performance is entirely unaccompanied, with no instrumental backing or background noise audible. The recording is dry and close-miked, with a subtle natural room reverb that enhances the vocal's presence without introducing digital artifacts.
reference caption (gemini 3 flash, the training target)This recording features a single male solo voice performing an intense Alternative Rock melody without any instrumental accompaniment or backing vocals. The singer demonstrates a wide dynamic range, starting with a breathy, intimate chest voice and building into a powerful, gritty belt in the upper-tenor register. His technique is marked by soulful legato phrasing, punctuated by strong glottal attacks and a natural, medium-width vibrato. The timbre is warm and resonant, yet possesses a raw, gravelly edge that enhances the defiant and determined emotional arc of the performance. The recording is dry and intimate, with a close-miked character that reveals subtle breath sounds and a hint of natural room reverb. The audio quality is clean and professional, devoid of any digital artifacts or background hum, focusing entirely on the expressive nuances of the unaccompanied voice.

Caption quality, measured on all 158 held-out clips

The gender table below scores one attribute. Writing a good caption is the model's actual job, and a checkpoint can agree about gender while writing worse captions — so caption quality was measured as a separate axis, and it is the axis that chose the default.

Method. Every checkpoint captioned all 158 held-out clips (greedy, 440 new tokens — the most Whisper's 448-position decoder allows). Gemini 3.6 Flash was then handed the audio and one caption and asked to score that caption, with every level of every axis anchored in words rather than left as a bare 0–5 scale. The judge is blind: it never learns which checkpoint wrote the caption, never sees a second caption and never sees the reference. It must write a one-line “what I hear” and list the statements it believes are wrong before scoring, so a surprising number is diagnosable. Every clip was judged twice and the passes averaged.

checkpointdetailfactual overalls.e.m.errors flagged clean captions
base — laion/music-whisper3.622.572.48 ±1.380.112.4215%
v1 frozen-encoder (previous release)4.873.213.23 ±1.350.111.6523%
v2 balanced, lr 1e-44.863.273.29 ±1.440.111.7129%
v2 unbalanced, lr 3e-5default4.913.453.47 ±1.410.111.4134%

Paired comparisons — every checkpoint captioned the same clips, so the question is not whose mean is higher but whether it wins on the same clip, against a 95 % bootstrap interval over clips.

comparisonoverall Δ95% CI pwon / tied / lostverdict
the two v2 contenders+0.18[-0.04, +0.39]0.0964 / 43 / 51within noise
default vs v1+0.24[+0.02, +0.47]0.0463 / 39 / 56separated
default vs base+0.99[+0.69, +1.28]<0.0196 / 18 / 44separated

The two v2 contenders are a tie on caption quality, and this page does not pretend otherwise. The gap between them is smaller than the judge's own disagreement with itself: re-judging the same caption a second time changes the overall score by 0.56 points on average, and the two passes agree exactly only 63 % of the time. The default goes to the unbalanced run because it is ahead on every other measured axis and behind on none, not because of 0.18 of a point here.

Fine-tuning does help captioning — by about three-quarters of a point over the base model, comfortably separated. That is worth saying next to the gender table, where the base model beats every fine-tune. The two axes genuinely disagree, and a page that reported only one of them would be selling a different model.

Detail is saturated; factual correctness is the whole story. All three fine-tunes sit at 4.86–4.91 for detail, with 135 of 158 clips at the maximum. They have all learned to write richly specific prose. What still separates them is whether the specifics are true.

The judge is an instrument, not ground truth: it is one model listening to synthetic singing, and on some clips it hears an instrument the reference caption says is not there. That bias falls identically on all four checkpoints over identical audio, which is what makes the paired differences meaningful and the absolute values only indicative. Raw judgements and the summary are published in the model repo under eval/.

Voice gender, measured on all 158 held-out clips

The gender problem was originally reported on ten clips. Ten is enough to notice a regression and far too few to choose a checkpoint on, so it was re-measured over the whole held-out split — 34 female and 124 male clips by the Gemini judge. Balanced accuracy is the mean of the two recalls, which is the number that matters here: plain accuracy is flattered by simply calling everything male on a split that is 78 % male.

checkpointfemale recallmale recall balanced acc.abstainsmean words
base73%73%73.1%1%135
v1 frozen-encoder9%99%54.0%0%137
v2 unbalanced lr3e-527%96%61.6%1%140
v2 unbalanced lr1e-421%96%58.3%0%139
v2 balanced lr1e-424%98%60.6%0%137
v2 balanced lr3e-512%99%55.5%0%141

The rebalancing did not work, and the controls are how we know. The two unbalanced rows are trained on the original corpus on the same 10-epoch schedule as the balanced ones. The best balanced accuracy of any fine-tune is an unbalanced run (61.6 %), ahead of both balanced runs (60.6 % and 55.5 %); at a matched learning rate the balanced run is worse than its control. So the improvement over v1's 54.0 % comes from the longer schedule with best-epoch selection, not from balancing the data.

Why balancing the audio changed so little. The corpus was balanced on the audio judge's verdict, but what the model is trained to reproduce is the caption text — and on the balanced set 2,342 of its 3,189 female-voiced clips (73 %) carry a caption that calls the singer male. The caption-text female share therefore moved only 5.6 % → 15.4 %, and female recall moved by almost exactly the same factor. The model tracked the label it was actually trained on. The next attempt has to fix the captions, not the audio mix.

The base model still wins this axis outright (73.1 %). If voice gender matters for your use, use laion/music-whisper.

Gender is read out of the generated caption by a conservative token-majority rule; captions that never commit are counted as abstains and excluded from the recalls, so a model cannot score well by saying less. The reference label is the Gemini audio judge — an instrument, not ground truth: a checkpoint can agree with it for the wrong reason, which is why validation loss is reported next to it and neither is used alone.

How to read this

Links

Predictions generated with greedy decoding, 440 max new tokens — the ceiling Whisper's 448-position decoder allows once its three forced start tokens are counted — on clips no checkpoint saw during training. These are the exact captions the caption-quality judge scored. An earlier revision of this page used a 200-token cap, which truncated between 15 % and 42 % of captions mid-sentence depending on the checkpoint; the voice-gender numbers are unchanged by the longer generation.