Speech Emotion Recognition on Open-Source Orukeet
Extended an open-source 25-language speech recognizer to hear how something is said, not just what is said: 16 emotions, delivery scales and word-level transcripts from one audio clip, served on GPU in about 0.2 seconds.
- Company
- FYI.AI
- Role
- ML engineering
- Period
- 2026
01 / Overview
Built a production speech-understanding system for FYI on top of Orukeet, an open-source 25-language speech recognizer. It was extended to hear how something is said as well as what is said. From one audio clip, the system returns:
- [01]
Emotions: 16 fine-grained labels, such as happy, excited, amused, frustrated, concerned and sad, each with a confidence score.
- [02]
Delivery: Continuous 0–1 scales for energy, positivity and dominance.
- [03]
Transcription: Word-level timestamps across 25 European languages.
02 / What was done
- [01]
Data: Assembled roughly 460 hours of emotional speech from more than 50 public corpora in about 20 languages. It covers acted speech, conversational dialogue, podcasts and synthetic voices. Every source went through automated quality checks before training (label consistency, duplicate detection and audio integrity).
- [02]
Modelling: Fine-tuned the Orukeet speech encoder to predict emotions and delivery together. Several generations of models were trained, and the final production model combines the strengths of two of them.
- [03]
Evaluation: Every result is measured on speakers the models never heard in training, on standard public benchmarks. Leakage between training and test data was audited, and the fixes were verified.
- [04]
Deployment: Served on GPU through an API. Full analysis of a 10–15 second clip (transcript plus emotions) takes about 0.2 seconds.
- [05]
Product: Shipped a dashboard workspace where users upload audio, choose a task and model, and see colour-coded emotion scores, delivery scales, the transcript, and timing details.
03 / Benchmarks
| Model | Emotion benchmark, 5,000 unseen clips (accuracy / F1) | Emotion benchmark, full set (accuracy / F1) | IEMOCAP positive / neutral / negative | IEMOCAP 4-class UAR | CREMA-D 6-class accuracy |
|---|---|---|---|---|---|
| Orukeet v4 (FYI, production) | 76.6% / 0.769 | — | 73.9% | 70.1 | — |
| Orukeet v3.1 (FYI) | 77.4% / 0.776 | 71.8% / 0.730 ¹ | 72.0% | 68.8 | 74.3% ³ |
| Orukeet v2 (FYI) | 53.3% / 0.505 | — ² | 74.4% | 71.3 | 79.6% ³ |
| emotion2vec+ large | 73.6% / 0.711 | 68.6% / 0.677 | — | — | — |
| emotion2vec+ seed | — | 68.7% / 0.680 | — | — | — |
| emotion2vec+ base | — | 68.5% / 0.683 | — | — | — |
| Oruk Resonance 1 / Spectra 1 | — | 77.8% / 0.816 ⁴ (58.4% on Oruk's unseen audio) | 71.5% | 66.3 | — |
| Oruk Resonance-2 | — | — | — | — | 61.7% |
| Oruk Fourier | — | — | — | — | 55.7% |
- FYI models are measured on speakers never heard in training. In the full-set, IEMOCAP and CREMA-D columns, the other models' numbers are their published results. In the 5,000-clip column, every model was run on the same clips.
- ¹ A 62,809-clip reconstruction of the benchmark (same sources and per-source counts), with speakers held out.
- ² Not measurable: v2 trained on speakers in that test set.
- ³ Held-out actors only.
- ⁴ Oruk trained on the benchmark's own sources and scored on its own clips.
04 / Results
- [01]
Beat emotion2vec+: The strongest open-source baseline, on identical unseen clips: 77.4% vs 73.6% accuracy, and 0.776 vs 0.711 macro-F1.
- [02]
Beat Oruk's published results: On IEMOCAP (74.4% vs 71.5%) and CREMA-D (79.6% vs 61.7%).
- [03]
Came within 0.4 points of Oruk's headline accuracy: On its benchmark (77.4% vs 77.8%), while being scored on speakers never heard in training.
- [04]
Shipped: GPU service, API and dashboard UI.