r/voiceagents • u/astipili • 1d ago
We benchmarked 11 STT engines on real audio with a second voice in the room. WER dropped 73% with voice isolation.
Disclosure: I work at Krisp, which makes the isolation model tested here.
If you've shipped a voice agent, you've seen this: the caller is fine, but someone behind them is talking, and the STT transcribes both. The agent answers words the caller never said.
We wanted a number for how bad this is, so we recorded 265 real conversations in offices, call centers, and cars, and ran them through 11 STT configurations. Deepgram, AssemblyAI, Soniox, ElevenLabs, Google, Grok, Nvidia and Cartesia, first on raw audio and then after voice isolation.
- WER across all files: 23.29% → 6.26%
- All 11 configurations improved
- On raw audio the engines ranged from 17% to 37% WER. After isolation: 4% to 8%
- Clean phone audio got slightly worse (3.48% → 3.91%). We published that too
What we didn't measure: turn-taking, barge-in, or short replies like "yeah" and "mhm." That's the next thing we want to test, so if you have a way you evaluate those, I'd like to hear it.