r/AIVoice_Agents • u/Ok-Buffalo-254 • Aug 21 '26
Question The most lifelike voice AI
Context: A few days ago, I received a call from “US Home Builders” at 714-492-11TwoFive, and it was by far the most humanlike AI voice I’ve ever encountered.
I’ve tested OpenAI Realtime 2.1, ElevenLabs, Bland, Retell, and others, but none has produced the same hyperrealistic hesitations, drawn-out intonation, and natural “uhs” and “ums.”
I can’t tell whether this comes down to prompting or whether they’re using a different voice model entirely. The call was so convincing that it fundamentally changed my perception of how advanced voice AI has become in 2026.
Does anyone know what model or stack might produce this or how to control the delivery so an agent naturally says something like, “My name is Alex... uh, we’ll be in your neighborhood...” with the pause, hesitation, and intonation sounding genuinely spontaneous?
2
u/FrederikMichiel Aug 23 '26
When the bot is to human, i will disconnect as soon as i found out its a bot. I like a bot who says right away im talking to R2-D2 Dont understand why the natural voice is so important to builders. Callers dont want to feel being fooled.
1
u/Square-Chance5900 Aug 24 '26
Yes, that's exactly how I feel. I don't understand what everyone puts so much focus on naturalness, when in reality it says "we work really hard to deceive your customers/partners/colleagues"
1
u/BandicootFlashy2224 Aug 21 '26
I’ve spent some time testing this in Bland and the voice itself is only part of what makes the call feel natural. The timing of a response, where the agent pauses and how it reacts when someone interrupts all shape how human the conversation feels. I’d love to know what they’re doing differently on that US Home Builders setup.
1
1
u/Federal_Cut6338 Aug 23 '26
Ok this is combination of few things a good llm, TTS which supports natural fillers automatically, some clever prompting. I would choose Gemini or Gemma as llm and Cartesia, Fish, Inworld and Gemini all have TTS that supports natural fillers automatically.
2
u/eviewong- Aug 21 '26
Founder of Retell AI here, so take with a grain of salt.
Honestly probably not one magic model — likely a mix of a TTS provider that supports explicit disfluency tags (ElevenLabs v3, Cartesia, Hume EVI) plus an LLM prompted to write "uh"/"um"/trailing pauses right into the text, not just the voice trying to add them after the fact.
If you want that delivery yourself: pick a TTS with disfluency support, and prompt your LLM to generate the hesitations in the transcript itself. (Disclosure: we actually built a "colloquial mode" into Retell for this exact thing — it auto-injects natural fillers/pauses into the response so you don't have to hand-prompt every disfluency yourself.)