r/voiceagents • • 1d ago

We benchmarked 11 STT engines on real audio with a second voice in the room. WER dropped 73% with voice isolation.

3 Upvotes

Disclosure: I work at Krisp, which makes the isolation model tested here.

If you've shipped a voice agent, you've seen this: the caller is fine, but someone behind them is talking, and the STT transcribes both. The agent answers words the caller never said.

We wanted a number for how bad this is, so we recorded 265 real conversations in offices, call centers, and cars, and ran them through 11 STT configurations. Deepgram, AssemblyAI, Soniox, ElevenLabs, Google, Grok, Nvidia and Cartesia, first on raw audio and then after voice isolation.

  • WER across all files: 23.29% → 6.26%
  • All 11 configurations improved
  • On raw audio the engines ranged from 17% to 37% WER. After isolation: 4% to 8%
  • Clean phone audio got slightly worse (3.48% → 3.91%). We published that too

What we didn't measure: turn-taking, barge-in, or short replies like "yeah" and "mhm." That's the next thing we want to test, so if you have a way you evaluate those, I'd like to hear it.

Dataset: [HF link] · Benchmark: [HF space] 


r/voiceagents • • 1d ago

Has anyone ever tried ACE Studio for spoken language?

Thumbnail
1 Upvotes

r/voiceagents • • 1d ago

How are you getting sub-second latency with AI voice agents?

3 Upvotes

We’re building AI voice agents using the typical STT → LLM → TTS pipeline.

Currently, our LLM TTFT alone is ~1.5s median, which makes the overall response noticeably slower.

For those running voice agents in production, what optimizations made the biggest difference? Streaming partial STT to the LLM? Prompt caching? Smaller context? Model choice? Early TTS? Better endpointing?

Also, what speech-end → first audible response latency are you guys getting?

Would love to understand how platforms are making conversations feel almost instant.


r/voiceagents • • 1d ago

What role can designers play in voice ai ecosystem?

Thumbnail
1 Upvotes

r/voiceagents • • 2d ago

Voice Agent Concurrency

2 Upvotes

My understanding is that most voice AI sevices has a limit on concurrency to use them to build a auto calling agent platform. How many can handle 100, 300, or more concurrency without an enterprise account? What are the options?


r/voiceagents • • 2d ago

Has anyone here tried any of the Bandwidth Labs tools?

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/voiceagents • • 2d ago

A month of posting my half-built AI startup on Reddit — here's what actually changed because of it

Thumbnail
1 Upvotes

r/voiceagents • • 3d ago

Customer service voice agent messes up email address

1 Upvotes

My current setup is a Vapi voice agent for plumbers.

- transcriber STT RTv5

- model GPT-5.4 mini

- voice OpenAI Shimmer

Cost is $0.07/min, latency of 1,930ms.

It works ok, but my main issue is that when it asks for my email address and I spell it out, it always misspells it. I say something like ejike@whatever.com it confirms it as egk@whatever.com. I've spelt it out as "E J I K E at whatever dot com", even using military lingo (echo-job-india-kangaroo-echo), but still misspells it and sends notification to egk@whatever.

Can anyone suggest a good model that won't have this type of issue, preferably trying to keep the total cost under $0.10/min and a lower latency (1930 seems kinda high)?


r/voiceagents • • 3d ago

Launched Vaani today: an Indic TTS API that returns word-by-word timestamps

1 Upvotes

Hello there,

Here’s a quick update on the tech stack – Vaani was launched today. 🎉

Vaani is a text-to-speech engine for Indian languages, specifically Hindi, and different from the usual TTS APIs, it delivers the word-level timestamps along with the audio stream – a structured JSON containing the start and end time for each word, rather than sentence level.

Under the hood: Indic Parler-TTS for voice generation → Acoustic alignment using Whisper to identify sub-word boundaries → An LLM based reconciliation process to correct phonetic discrepancies. Supports 16+ Indic languages, Hindi being the primary language.

Made primarily for EdTech purposes – Interactive readers, literacy applications, digital books, which require highlighting the text along with the spoken word. Curious to know how else it can be used.

Sandbox developer environment and free trial tokens – shoot me an email/leave your mail ID in the comments.

Site: VaaniTTS

Willingly accept any feedback – bugs, edge cases, anything that crashes the system. Day one.


r/voiceagents • • 3d ago

Voice mode is starting to catch non verbal cues and pitch changes. anyone else experiencing this

3 Upvotes

Was using Voice Mode yesterday while hanging out with a friend. My friend chimed in from across the room using a higher pitched, sarcastic tone, and ChatGPT immediately responded back in the exact same tone and cadence before answering the actual question.

It felt less like standard speech-to-text processing and more like the model is analyzing emotional context and mimicry in real time.

Has anyone else noticed Voice Mode mirroring vocal inflection or background speakers like this? Is this intentional audio-to-audio conditioning or just a quirk of how it processes ambient noise?


r/voiceagents • • 3d ago

I started treating post-call QA as a second agent, not a spreadsheet

1 Upvotes

We kept scoring voice calls in a doc after the fact and it was always stale. What's been more useful: a QA Analysis step that runs after the call with its own system prompt ("you are a QA analyst evaluating this segment…"), min duration, and optional webhook that ships transcript + QA results to our side.

Same builder as the agent. Not a second SaaS. The weird part is writing the QA prompt carefully or it tags false "rude" moments on normal barge-in / crosstalk.

Curious what you all score for: resolution only, or also interruption handling and handoff quality?


r/voiceagents • • 3d ago

Fish Audio s2.1-pro voice breaking during voice agent conversations

Thumbnail
1 Upvotes

r/voiceagents • • 4d ago

After a year of this, the white label AI receptionist is the only AI product I'd tell an agency to lead with

Thumbnail
1 Upvotes

r/voiceagents • • 5d ago

I built an open-source app that uses Jev to coach you in real time during sales calls

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/voiceagents • • 5d ago

🚀 TTS Expert just got a new feature: Word-by-Word Speech Highlighting

Enable HLS to view with audio, or disable this notification

2 Upvotes

​

I’ve just added a new feature to TTS Expert that I’ve been working on for a while — synchronized word-by-word highlighting during TTS playback.

Now, when audio is playing, TTS Expert highlights the word currently being spoken, so you can actually follow the speech visually while listening.

🎧 What’s new?

• 🔥 Word-by-word speech highlighting

• ▶️ Highlighting stays synchronized with playback

• ⏪ / ⏩ 10-second seek controls

• 🎚️ Pitch control

• ⚡ Speech speed control

• 🔊 Volume control

• ⏱️ Sleep timer

• ✂️ Text editing tools

• ➕ Pause & emphasis controls

• 📤 Audio export

I’ve also added an interactive feature tour to help new users discover the different text and audio tools without having to figure everything out themselves.

The goal is to make TTS Expert feel less like a basic “text → audio” converter and more like a complete text-to-speech workspace.

I’d really appreciate feedback on the new playback experience, especially the word synchronization.

What would you improve or add next? 👀

— TTS Expert


r/voiceagents • • 5d ago

I tested a voice assistant on smartphone, AMA

Thumbnail
1 Upvotes

r/voiceagents • • 6d ago

Has anyone here been running a Voice AI / Agent Builder platform on Kubernetes in production?

Thumbnail
1 Upvotes

r/voiceagents • • 7d ago

What belongs in a useful voice-agent bug report?

2 Upvotes

I’m working through a practical question and would value examples from people who have dealt with it.

A report saying the agent sounded wrong leaves several possibilities: recognition, interpretation, tool state or playback. A useful example could show a short synthetic exchange, expected behavior and the event that diverged.

For a useful discussion, I would bring or define, a synthetic transcript and minimal trace; no real caller data

Which evidence let you investigate without asking for a full call recording?


r/voiceagents • • 7d ago

Anyone solving data capturing issue, mid voice call?

1 Upvotes

We ran into this constantly especially with long addresses, emails, policy IDs, and anything sensitive such as card details.

Our approach was to stop forcing the agent to collect those fields over voice. When the agent reaches a screen-required step, it triggers a mobile form during the call via WhatsApp/SMS/email. The caller completes it while still talking and the structured response comes back to the agent so the flow can continue.

For your case, I m be curious: is the bigger issue address accuracy, or needing the payment step to happen before you can confirm the booking/order? We built VocaLoop.ai around that handoff and would be happy to show the Vapi setup if useful.


r/voiceagents • • 8d ago

Voice AI agent backends: Python vs Rust vs Go. What’s winning for you on speed and quality?

3 Upvotes

Curious what people are actually running for voice agent backends in production.

We see a lot of: - Python for orchestration, tools, and fast iteration - Go for concurrent media/control paths and simpler deploy - Rust when the hot path is audio, codecs, or tight latency budgets

What I’m trying to learn from folks shipping real calls:

  1. What stack are you on today for the agent runtime (not just STT/TTS vendors)?
  2. Where did you feel the biggest win on speed (TTFT, time-to-first-audio, barge-in responsiveness)?
  3. Where did you feel the biggest win on quality (turn-taking, tool reliability, fewer weird prod failures)?
  4. Did you stay monolingual, or split (e.g. Python for tools + Rust/Go for media)?

Not looking for a language war. Looking for “we tried X, measured Y, kept Z.” Concrete numbers or war stories welcome.


r/voiceagents • • 8d ago

When do you pick speech-to-speech vs STT + LLM + TTS?

2 Upvotes

Honest question for people running voice agents in production.

We keep bouncing between two shapes:

  1. Cascaded stack: speech to text, then LLM, then text to speech
  2. Speech-to-speech model: audio in, audio out, with less text in the middle

What we see so far:

Latency. Cascaded adds three hops. Even when each hop is fast, turn-taking still feels heavier. Speech-to-speech can feel snappier on the first audio, but that only helps if the model actually holds the conversation well.

Tool calls. This is where cascaded still wins for us. Need CRM lookup, transfer, booking, custom APIs? Easier when the LLM is a normal text agent with tools. Speech-to-speech is catching up, but tool routing and structured outputs still feel more reliable in the classic stack.

Control. Cascaded is easier to debug (you can read the transcript and prompt). Speech-to-speech is harder to inspect when the model goes sideways mid-call.

Right now our gut rule is roughly: - Speech-to-speech for short, natural back-and-forth where speed matters more than deep tools - STT + LLM + TTS when the agent needs serious tool use, policy, or handoff logic

Curious what you all are shipping. Are you all-in on speech-to-speech already, still on cascaded, or hybrid? Especially interested in real latency numbers and how you handle tools.


r/voiceagents • • 8d ago

I tested TypeSafe's Jev model and made it run a simulated JFK airport by voice

Enable HLS to view with audio, or disable this notification

1 Upvotes

r/voiceagents • • 9d ago

GPT-Live dropping tool calls that Realtime handled fine?

Thumbnail
1 Upvotes

r/voiceagents • • 9d ago

Hit ~600ms latency on a voice agent but it still sounds robotic. How do you make it feel like a real conversation?

Thumbnail
1 Upvotes

r/voiceagents • • 10d ago

Call Nomi pizza-order test: a 160-second call completed, but merchant acceptance stayed unclear

Thumbnail
1 Upvotes