Skip to content
All work

Teaching a Voice Model to Cough, Sigh and Laugh

Fine-tuned the open-source OmniVoice text-to-speech model to perform non-verbal sounds on cue, like a cough or a sigh written inline in the script, and shipped it in FYI's voice service.

Company
FYI.AI
Role
ML engineering
Period
2026
PyTorchHugging Face TransformersOmniVoiceWhisperDockerGitHub Actions

01 / At a glance

8New expression tags, each its own token in the model
111 hTraining audio from 7 speech corpora
0Tag words read out loud across 180 test clips
6Training rounds to the shipped model

02 / The problem

How alike the cough and laughter tags sounded
  • Base model99.7%
  • Training round 198.2%
  • Training round 398.0%

Acoustic similarity, where 100% means identical. Asking for a cough produced a laugh, and neither 19× more cough data nor 25 epochs pulled them apart.

03 / Why it happened

Tag in the scriptPieces the model actually saw
[cough]12447 · 1384 · 60
[laughter]58 · 51783 · 60
[clears throat]58 · 7422 · 82 · 27591 · 60
  • Tags were not symbols to the model, just ordinary text split into sub-word pieces. They shared the same bracket pieces (58, 60), so the rarer tag blurred into the commoner one.

04 / The fix

  1. 01Before
    • [cough] = 3 shared pieces
    • Blurs into laughter
  2. 02After
    • <|cough|> = 1 token
    • Its own embedding
    • Vocabulary 151,676 → 151,684

05 / How it was built

  1. 01Diagnose
    • Tags split into shared pieces
    • Cough ≈ laughter
  2. 02New tokens
    • 8 dedicated tokens
    • One embedding each
  3. 03Data
    • 7 corpora
    • Emoji labels → inline tags
    • 28% plain speech
  4. 04Train
    • 6 rounds
    • 15 epochs
    • 29,730 steps
  5. 05Ship
    • FYI voice API
    • One-step rollback

06 / The data hunt

~950Usable English cough clips in every open dataset combined
243 of 1,116Meeting-corpus coughs with speech around them
24,991TED-LIUM "cough" labels, and not one audible cough

07 / Where the data came from

SourceClipsWhat it added
NonVerbalSpeech-38K30,509Cough, throat clear, sigh, laughter, sniff, breath
LibriTTS-R8,133Plain speech, so everyday voice quality holds
SMIIP-NV6,343Studio coughs in a neutral register
NonverbalTTS5,045Breath, laughter, English cough
MNV-171,455The only sneezes, plus every other tag
AMI Meeting Corpus243Real English coughs in conversation
  • Rule: a tag must sit at the exact word where the sound happens. Clip-level labels and spliced-in sound effects were excluded.

08 / Training data by tag

Share of tagged training clips
  • Cough26.8%
  • Sigh18.8%
  • Laughter16.1%
  • Breath14.0%
  • Sniff13.6%
  • Clears throat9.7%
  • Sneeze0.4%

Out of 45,196 tagged clips. Sneeze had only 178 examples and still passed listening tests. Groan (248 examples) was trained but held back from the API.

09 / More data did not help

Tags that fired, by checkpoint
  • Round 5 (previous release)51%
  • Round 6, epoch 3.846%
  • Round 6, epoch 7.553%
  • Round 6, epoch 11.2 (shipped)50%
  • Round 6, epoch 1548%

Round 6 had 2.3× the cough data from three new corpora. Every checkpoint sits within the ±6-point measurement error, so the answer was not another training round.

10 / What actually moved it

Tags that fired, by reference voice
  • Voice A, relaxed and spontaneous84%
  • Voice B58%
  • Voice C48%
  • Voice D, formal read10%

Same model and settings, every reference trimmed to about 7 seconds. The voice being cloned mattered more than any checkpoint: an 8× spread.

11 / How it was measured

  1. 01Transcript
    • Tag word never spoken
    • 0 of 180 clips
  2. 02Tag detector
    • Built for this project
    • Checks for a sound at each tag
    • 1% false alarms
  3. 03Test loss
    • Repeatable scoring
    • Tagged and plain speech
  4. 04Listening
    • Final say on every release

12 / Fixes after launch

6.0% → 3.9%Dead air after storing word-for-word voice transcripts
+1.3 kHzMore audio detail at the same file size
2 of 4Takes that lost the laugh on the new token, so laughter went back to the model’s own
0.99Voice similarity to the reference speaker

13 / Lessons

  • [01]

    Check the representation before adding data: 19× more cough data changed nothing until the tags became real tokens.

  • [02]

    Counting events is not judging quality: The detector preferred settings that sounded worse. Listening overruled it four times.

  • [03]

    The biggest wins were outside the model: Voice choice, transcripts and audio encoding beat every extra training round.