Skip to content
All work

Speech Emotion Recognition on Open-Source Orukeet

Taught an open-source speech recognizer to hear how something is said, not just what is said.

Company
FYI.AI
Role
ML engineering
Period
2026
PyTorchHugging Face TransformersFastAPIDockerNVIDIA H100Next.js

01 / At a glance

77.4%Accuracy on speakers never heard in training
+3.8 ptsAhead of emotion2vec+, the best open-source baseline
0.2 sTo analyse a 10–15 s clip on GPU
460 hEmotional speech from 50+ public corpora

02 / One clip in, three answers out

16Emotions, each with a confidence score
3Delivery scales: energy, positivity, dominance
25Languages transcribed with word timestamps

03 / How it was built

  1. 01Data
    • 460 h of speech
    • 50+ corpora
    • ~20 languages
    • Automated quality checks
  2. 02Train
    • Orukeet encoder
    • Emotion + delivery together
    • Best two generations merged
  3. 03Evaluate
    • Unseen speakers only
    • Public benchmarks
    • Leakage audited
  4. 04Ship
    • GPU API
    • 0.2 s per clip
    • Dashboard UI

04 / Benchmarks

This project Other models
Emotion benchmark · 5,000 unseen clips
  • Orukeet v3.177.4%
  • Orukeet v4 (production)76.6%
  • emotion2vec+ large73.6%
IEMOCAP · positive / neutral / negative
  • Orukeet v274.4%
  • Orukeet v4 (production)73.9%
  • Orukeet v3.172.0%
  • Oruk Resonance 171.5%
CREMA-D · 6 emotions, held-out actors
  • Orukeet v279.6%
  • Orukeet v3.174.3%
  • Oruk Resonance-261.7%
  • Oruk Fourier55.7%

Accuracy, %. Orukeet models are scored on speakers never heard in training. In the 5,000-clip test every model ran on the same clips; the other IEMOCAP and CREMA-D numbers are the vendors' published results.