Teaching a Voice Model to Cough, Sigh and Laugh
Fine-tuned the open-source OmniVoice text-to-speech model to perform non-verbal sounds on cue, like a cough or a sigh written inline in the script, and shipped it in FYI's voice service.
- Company
- FYI.AI
- Role
- ML engineering
- Period
- 2026
01 / At a glance
02 / The problem
- Base model99.7%
- Training round 198.2%
- Training round 398.0%
Acoustic similarity, where 100% means identical. Asking for a cough produced a laugh, and neither 19× more cough data nor 25 epochs pulled them apart.
03 / Why it happened
| Tag in the script | Pieces the model actually saw |
|---|---|
| [cough] | 12447 · 1384 · 60 |
| [laughter] | 58 · 51783 · 60 |
| [clears throat] | 58 · 7422 · 82 · 27591 · 60 |
- Tags were not symbols to the model, just ordinary text split into sub-word pieces. They shared the same bracket pieces (58, 60), so the rarer tag blurred into the commoner one.
04 / The fix
- 01Before
- [cough] = 3 shared pieces
- Blurs into laughter
- 02After
- <|cough|> = 1 token
- Its own embedding
- Vocabulary 151,676 → 151,684
05 / How it was built
- 01Diagnose
- Tags split into shared pieces
- Cough ≈ laughter
- 02New tokens
- 8 dedicated tokens
- One embedding each
- 03Data
- 7 corpora
- Emoji labels → inline tags
- 28% plain speech
- 04Train
- 6 rounds
- 15 epochs
- 29,730 steps
- 05Ship
- FYI voice API
- One-step rollback
06 / The data hunt
07 / Where the data came from
| Source | Clips | What it added |
|---|---|---|
| NonVerbalSpeech-38K | 30,509 | Cough, throat clear, sigh, laughter, sniff, breath |
| LibriTTS-R | 8,133 | Plain speech, so everyday voice quality holds |
| SMIIP-NV | 6,343 | Studio coughs in a neutral register |
| NonverbalTTS | 5,045 | Breath, laughter, English cough |
| MNV-17 | 1,455 | The only sneezes, plus every other tag |
| AMI Meeting Corpus | 243 | Real English coughs in conversation |
- Rule: a tag must sit at the exact word where the sound happens. Clip-level labels and spliced-in sound effects were excluded.
08 / Training data by tag
- Cough26.8%
- Sigh18.8%
- Laughter16.1%
- Breath14.0%
- Sniff13.6%
- Clears throat9.7%
- Sneeze0.4%
Out of 45,196 tagged clips. Sneeze had only 178 examples and still passed listening tests. Groan (248 examples) was trained but held back from the API.
09 / More data did not help
- Round 5 (previous release)51%
- Round 6, epoch 3.846%
- Round 6, epoch 7.553%
- Round 6, epoch 11.2 (shipped)50%
- Round 6, epoch 1548%
Round 6 had 2.3× the cough data from three new corpora. Every checkpoint sits within the ±6-point measurement error, so the answer was not another training round.
10 / What actually moved it
- Voice A, relaxed and spontaneous84%
- Voice B58%
- Voice C48%
- Voice D, formal read10%
Same model and settings, every reference trimmed to about 7 seconds. The voice being cloned mattered more than any checkpoint: an 8× spread.
11 / How it was measured
- 01Transcript
- Tag word never spoken
- 0 of 180 clips
- 02Tag detector
- Built for this project
- Checks for a sound at each tag
- 1% false alarms
- 03Test loss
- Repeatable scoring
- Tagged and plain speech
- 04Listening
- Final say on every release
12 / Fixes after launch
13 / Lessons
- [01]
Check the representation before adding data: 19× more cough data changed nothing until the tags became real tokens.
- [02]
Counting events is not judging quality: The detector preferred settings that sounded worse. Listening overruled it four times.
- [03]
The biggest wins were outside the model: Voice choice, transcripts and audio encoding beat every extra training round.