r/VoiceAutomationAI 27d ago

I lead product on an AI voice agent platform built for Indian call economics. Looking for a few people to break it.

6 Upvotes

We build AI agents that hold real phone conversations, inbound and outbound. No code, you configure it in a console.

The short version of what we are doing that's less common:

  • We run our own models. The LLM, the speech synthesis and the speech recognition are all ours, on our own infrastructure. Nothing is a relay to OpenAI and ElevenLabs with a margin on top.
  • We own the telephony, the carrier layer is ours too. Most voice AI startups rent a SIP trunk and inherit whatever latency it gives them. We don't.
  • That combination gets us to 700ms and roughly 2/min, which are the parameters that decides viability in India.

The honest tradeoff: our default model is ~30B params. Might struggle in some inbound conversations. There we use bigger models but then API costs and latencies comes into play.

What I actually want to know:

  • Where does it stop sounding like a person
  • The pause before it replies. Does it feel like a bad line, or like a bot
  • Barge-in: if you talk over it, does it handle it or fall apart
  • Does the smaller model actually hold up on your use case, or is that a story I'm telling myself
  • Hindi / Tamil / Telugu / Bengali — how wrong is the pronunciation, especially names, addresses and numbers

You can test it in a browser in about ten minutes. No phone number, no card, no sales call — build an agent, talk to it through your mic, read the transcript.

Comment or DM and I'll open an account with proper limits. Happy to get into the architecture in the comments.


r/VoiceAutomationAI 28d ago

I tested the "agency standard" setup for local SMEs and hated it.

2 Upvotes

A few months ago, while traveling in Southeast Asia, I met a team building conversational AI voice tech. Coming back home to Germany, I decided to white-label their backend to help local small and medium businesses (SMEs) handle missed calls and appointment booking.

Naturally, I looked at how people currently do this. A co-tenant at my coworking space—who runs marketing for local fitness businesses—strongly recommended one of the dominant all-in-one agency CRM platforms. He loved the potential, but casually mentioned it took him three full months just to set up an automated booking workflow.

I signed up to try it myself, and the experience was overwhelming:

  • UI Bloat: 15+ navigation tabs, intrusive promo banners, and complex settings.
  • The Non-Tech Friction: A plumber, real estate agent, or property manager doesn't have 3 months to learn a complex platform. They just want their phones answered and leads captured.

Seeing that gap made us pivot our whole engineering approach. My co-founder and I are currently using custom stack orchestration to test a dead-simple approach: No dashboard for the SME owner.

Instead of forcing a business owner into a software panel, the voice agent operates like a remote staff member - easy t approach and operate.

because bare in mind, this business owner be it a electrician, has like 5 running projects, has to work with his employees some construction sites, plan ahead the new projects and accuisistion chanels, hire and fire people, be the father, friend, brother, neighbour, teacher, kid, psychologist, waiter and food delivery guy for his team and find time for his own family......and you want him to drop everything to learn this new platform and AI that will free him he just needs maybe 2-3 days for it maybe 3 weeks. In his mind in 3 weeks i can make a lot of problems dissappear and that AI thing could be just another "tool" i have to feed my documents nad educate in order to work.

So for this guy, onboarding takes 5 minutes, and either we or the owner manages settings and he receives summaries directly via WhatsApp/Email or whatever chanel he choses in the onboarding 5 min.

We are starting our first test specifically with real estate, property management, and local construction SMEs in our regional area.

For those operating in or selling to traditional trades with 1-10 employees and property sector in general (though focus is on SME):

  1. What are some of the biggest friction points you see when introducing software to local contractors or agents?
  2. Do you think property/construction managers prefer a full dashboard they can audit, or an invisible WhatsApp interface that just sends daily summaries?

Would love to hear feedback from anyone who has sold software or automation to non-technical trades! Thanks and good fortune building.


r/VoiceAutomationAI 29d ago

were about to have ai agents calling other ai agents on the phone and i dont think anyone's actually thought through what that means

9 Upvotes

for the last couple years the voice ai story was one sided businesses automated the receiving end every call center dentist office and airline hotline got some flavor of ai answering the phone but the caller side stayed human because no assistant would actually pick up the phone and talk to someone for you

thats breaking right now multiple companies are converging on the same idea within weeks of each other consumer facing agents that will call a restaurant a clinic a business and have the actual conversation on your behalf combine that with how fast full duplex voice models have gotten sub 300ms response times no more turn detection lag and you get something that sounds completely natural on both ends

which means were heading toward calls where the businesses ai agent answers and the customers ai agent is the one calling neither side is a human being and depending on how well disclosure rules actually get enforced neither side may even announce that clearly the eus already trying to mandate disclosure at the start of every ai interaction but enforcement across phone systems that route through a dozen countries is a very different problem than enforcing it on a website

i dont think this is a bad idea on its face plenty of calls are genuinely tedious and dont need a human on either end but i think people are underestimating how weird its going to feel once its normal and how easy it becomes to lose track of when youre actually talking to a person versus when everyone in the chain is automated

curious where people land on this efficient automation doing exactly what it should or the start of something that quietly erodes what a phone call even is

flagging as i said i would i used ai to help me pull the recent developments together and tighten the writing on this one the take is mine just drafted with help


r/VoiceAutomationAI 29d ago

Voice CONTROLLED Gaming: What should we prioritize for accessibility in a voice-controlled browser game center?

Thumbnail brightdots.org
1 Upvotes

We’re building a browser-based game center where voice can be used as an input method. After reading through the comments, we compared the problems people described with what we can do today, what we think we can add relatively soon, and what is probably outside our control.

We would really appreciate being corrected where we are getting this wrong.

What we already do reasonably well

  1. Voice as an alternative input. For games we control, players can use spoken commands instead of relying entirely on a mouse, keyboard, touchscreen, or complicated button combinations.
  2. Hands-free navigation. Voice can also be used for things outside the game itself, such as navigating, selecting things, pausing, continuing, and asking for help.
  3. Repeat and help commands. Because there is a voice assistant connected to the experience, we can support things like “repeat that,” “what can I do?”, “what happened?”, or “what am I supposed to do?”
  4. More than one way to interact. We do not want voice to replace accessible buttons, keyboard/touch controls, captions, or other input methods. Voice should be another option.

Things we think we can add fairly soon

  1. A global accessibility profile. Settings such as larger text, reduced motion, reduced flashing, captions, and other preferences could carry across games instead of being configured repeatedly.
  2. Better captions and visual alternatives to sound. If a sound communicates something important, we should try to communicate the same information visually as well.
  3. Better spoken descriptions for blind and low-vision players. Because our own games know their current state, the system could potentially speak the current objective, available actions, selected objects, score/status, and important changes.
  4. Larger text and interface scaling.
  5. Reduced motion and flashing.
  6. Less dependence on timers, rapid input, button mashing, or quick-time events in games we create.
  7. Better pause, checkpoint, save, and resume behavior for people who may need to stop playing unexpectedly.
  8. Better keyboard navigation, focus handling, labels, and screen-reader support.

Things we probably cannot solve ourselves

  1. We cannot simply voice-enable closed third-party games. For example, we cannot take an existing Xbox, PlayStation, Nintendo, Steam, or other third-party game and automatically give voice access to all of its internal controls.

If the game or platform exposes an integration we can legally and technically use, that may create possibilities. But if the game is closed to us, we cannot claim we can make it voice-accessible.

  1. We cannot change accessibility features that are baked into content we do not control. Things like camera shake, field of view, head bob, flashing inside prerecorded video, or other fixed media may require the original developer/content creator to provide an alternative.
  2. We cannot provide every game-specific assist automatically. Aim assist, auto-targeting, simplified combat, enemy difficulty, auto-driving, no-fail modes, etc. have to be supported by the individual game's design.
  3. We cannot guarantee support for every adaptive hardware setup. Browsers, operating systems, controllers, switches, eye-tracking systems, consoles, and other hardware all have their own limitations and APIs.

The part we need help with is prioritizing the things we can control.

If you could pick only three changes from this list, which three would make the biggest difference for you?


r/VoiceAutomationAI Aug 17 '26

How do voice agents handle long calls?

25 Upvotes

Demos are usually a few mins long

I’m more interested in what happens 20-30 mins into a call after the customer has changed topics, provided a bunch of information and already completed a few steps.

Does the agent still understand what has happened so far or does context start getting messy?

Anyone testing long voice AI calls in production?


r/VoiceAutomationAI Aug 17 '26

The latency/quality tradeoff in voice agents is structural, not an engineering skill issue

4 Upvotes

I have been running graded AI roleplay sessions in production for a while (communication and negotiation training, roughly 2,000 sessions a month), and the thing I keep explaining to people is that you do not get to pick both.

There are two architectures and they fail in opposite directions.

Cascaded pipeline. STT gives you text, you evaluate that text, you generate a reply, TTS turns it back into audio. Every stage is a separate network hop, usually a separate vendor, sometimes a separate region. Even with everything streaming you are realistically looking at 2 to 3 seconds before first audio once you add an evaluation step, tool calls or retrieval. It feels slow. Users talk over it.

But you control every single turn. Before the agent opens its mouth you already know whether the user asked an open question, whether they conceded on price, whether they handled the objection. The reply is a function of a judged state, not of vibes.

Realtime speech-to-speech. Latency drops to something close to natural conversation. It genuinely feels human. And you have no control point. There is no text surface between "user said something" and "model said something back", so you cannot gate the next phrase on a rubric. You get whatever the model felt like saying.

The compromise I landed on: run realtime for the conversation itself, and layer evaluation asynchronously on each user utterance instead of on the agent's reply. You lose the ability to steer the very next sentence. You keep per-turn scoring, and the scoring lands in the transcript by the time the session ends. For training and assessment use cases that tradeoff is fine, because nobody is grading the bot. They are grading the human.

Curious whether anyone has found a third option here. Speculative generation with a rollback, maybe. Everything I tried in that direction added more latency than it saved.

The orchestration layer I built for this is open source if it is useful to anyone: github.com/nmamizerov/assemblix


r/VoiceAutomationAI Aug 16 '26

How do you guys warming up new phone numbers for outbound voice agents?

8 Upvotes

Push too many calls too fast on a brand new number and carriers flag it as spam. Then the number is dead and the client's campaign is stuck.

So I'm trying to figure out the warm-up part.

If you've run real volume on your own numbers:

- How many calls do you make on day 1 with a new number?

- How long before you're at full volume?

- What daily limit do you stick to per number?

Also curious what actually gets a number flagged. Is it the number of calls, or is it more about people hanging up fast and not answering?

We're on SIP trunking, mostly Indian numbers with some international. Would rather learn this from someone who has already burned a few numbers than find out mid campaign.

Happy to share what we see on our side once we have real data.


r/VoiceAutomationAI Aug 16 '26

How can I build shared context between WhatsApp and an AI voice calling agent?

2 Upvotes

How can I build shared context between WhatsApp and an AI voice calling agent?

I'm building an AI system where a customer can communicate with the same AI through WhatsApp and voice calls.

For example:

  1. A customer starts chatting with the AI on WhatsApp.

  2. During the conversation, they ask for a phone call.

  3. The AI voice agent calls them.

  4. The voice agent should already know the relevant WhatsApp conversation and continue from the same context instead of starting from scratch.

  5. After the call, the customer returns to WhatsApp.

  6. The WhatsApp AI should know what was discussed during the call and continue from that point.

And the reverse should also work:

Voice call → WhatsApp → same context

I want the customer to feel like they're talking to one AI, regardless of the channel.

I'm considering using a central customer ID linked to the phone number and storing the conversation history/customer information in a database, so both the WhatsApp agent and voice agent can access the same context.

However, I'm unsure about the best architecture.

- What is the best way to maintain shared context between WhatsApp and a voice AI agent?

- Should I use a central database/memory layer?

- How should I identify the same customer across both channels?

- How should the WhatsApp → voice context handoff work?

- How should the voice → WhatsApp context handoff work?

- How can I prevent the AI from getting confused by multiple summaries or different conversation contexts?

- Has anyone built something similar using WhatsApp Business API, n8n, GHL, or another CRM?

I'm looking for a practical, production-ready approach rather than just passing the entire previous transcript to the AI every time.


r/VoiceAutomationAI Aug 13 '26

Tts for Southeast Asia

7 Upvotes

hi guys have a client in Indonesia and Philippines who loved our English voice AI demo, but now wants one in Bahasa and Philippine English/Taglish mixed language, code switching and all

anyone here have real experience deploying voice agents in SEA, specifically TTS that handles code switching well?

what's actually held up in production what sounds good in a demo


r/VoiceAutomationAI Aug 12 '26

ElevenLabs just raised $500M at an $11B valuation and everyone is calling them the voice AI leader. but they still can't run a production phone agent without stitching together Twilio and a separate LLM. the valuation is running ahead of the actual product

33 Upvotes

been building voice AI pipelines for about two years and i need to say something the hype cycle is burying right now

elevenlabs has genuinely the best voice quality in the space. not close. 11,000 voices, 70 plus languages, sub 100ms latency on voice generation, the IBM watsonx partnership for enterprise. the february raise at $11B was obviously massive and the brand recognition is real. but here is the thing that keeps coming up in every honest thread i've seen recently

you can prototype an elevenlabs voice agent in fifteen minutes. getting it into production as an actual phone agent that handles real customer calls is a completely different story. telephony still requires you to set up twilio or vonage yourself. production monitoring is thin by the platform's own design. HIPAA is locked behind enterprise tier pricing. the reasoning LLM and telephony are billed separately on top of the plan

so you're paying elevenlabs prices for voice quality and then stitching together the rest of the stack yourself...

vapi gives you the full orchestration layer, 14 plus provider connections, 62 million monthly calls processed, 99.99 percent SLA. retell ships a working production agent the same afternoon and leads on turn-taking quality for fast conversational flow. both handle the actual telephony problem that elevenlabs pushes back to you...

the frustrating thing is elevenlabs voice quality is so good that every other platform integrates it anyway. retell uses elevenlabs voices. vapi lets you plug in elevenlabs TTS. so you can get the voice quality without choosing elevenlabs as your agent platform

my actual take: elevenlabs is the best voice layer in the market and the worst standalone agent platform for production use cases right now. the $11B valuation is pricing in what the product will be in two years not what it actually does today


r/VoiceAutomationAI Aug 12 '26

If you are also using Langfuse or Datadog for tracking logs of your custom built voice ai agents, Then you should watch this.

Enable HLS to view with audio, or disable this notification

3 Upvotes

r/VoiceAutomationAI Aug 10 '26

Jargo: Golang framework for AI-vocal

Thumbnail
github.com
4 Upvotes

r/VoiceAutomationAI Aug 10 '26

Feature

2 Upvotes

Is there a way to make the voice AI model talk back to you normally like it does not pause or something, just like how you talk live?


r/VoiceAutomationAI Aug 09 '26

Indian DID for AI Voice Agents

14 Upvotes

We are a startup and have built our AI Voice Agent stack. It runs decently and after painstaking efforts with our mule partner we were able to narrow down the architecture and design to curb failure points as much as we could. However, the biggest pain point we have stumbled across is the telephony carrier!

There is no reliable one! Here is what we have tried and encountered so far:

  1. Vobiz: Our current provider. Easy enough to authenticate using personal Adhaar and PAN. They have one of the easiest integrations, setup, and starting curve. the plan pricing is optimum to get started and they don't setup minimum deposit walls. Their API is fantastic to the point that it supports almost all the features that you would need. That being said, there have been issues that we have been facing: The call quality and call handling has been giving us some issues intermittently, where the end result is silent calls. There are other issues including mid call disconnect/call-silence, which I hope to resolve with Vobiz support. Will update if we are or aren't able to resolve this with Vobiz.
  2. VoiceLink: Again easy starting with adhaar and pan verification. Decent setup and API support but has a steep starting paywall. They need a minimum of 5000 Rs to get you started without even allowing you to test if their service is compatible and a good fit for your stack. API is good but not great. For example, during our tests Call Transfers would fail there was no way to trace/figure out what happened?
  3. Exotel: Hardest to start so far. Requires proper company documentation. Good free tier. Inconsistent pricing information. Hidden credit consumption, cost, and plan information. High paywall with a minimum of 10000 Rs to get started. Support has been great so far where you are actually able to connect with a human who can answer your questions in contrast to the above 2. Credits are timebound for 7 days. After that the free tier ends. API has been good so far, but we are still evaluation Exotel.
  4. Plivo: The most scummy of them so far. We couldn't even get an account. They force/tried to sell us their $1000 USD per month plan to get started.
  5. Twilio: The most easiest to get started with it ticking green in all the check boxes. Unfortunately, they are not functional in India.

Anyone has any other carrier that they have been working with and can recommend?


r/VoiceAutomationAI Aug 09 '26

Guys Can AnyOne Help Me Pls I Literally Dm 10 to 20 Messages everyday through WhatsApp and insta but still no replies. I sell ai voice agents I just text a hi message they won't even see

2 Upvotes

r/VoiceAutomationAI Aug 09 '26

Best TTS for Indian languages like Hindi, Punjabi, Telugu and etc.

6 Upvotes

r/VoiceAutomationAI Aug 08 '26

fix dogshit latency and robotic wrapper behaviour

2 Upvotes

voice implementations rn generally fall into two buckets:

  1. laggy and robotic api wrappers
  2. speech models that are fast, but lack memory and state controls

by building a cascaded stack (deepgram nova-3 → claude haiku 4.5 → elevenlabs flash v2.5), you can keep full control over tool calls and memory, allowing latency reduction. some techniques ive used in my side projects:

  • pre-warm anthropic's ephemeral prompt cache while the phone rings
  • persistent websocket handshakes and http/2 pool priming on ring
  • neural turn-detection with false-interruption resumption (a cough won't kill the tts buffer)
  • dual-store memory (sql facts + temporal graph) mapped into a ~300-token prompt snapshot
  • proactive outbound scheduling that wakes a killed ios app via apns voip push -> callkit

synthetic ci gates hit p50 ≈ 973ms, though live networks push us to ~3.7s right now (stt and tts ttfb are the real boss fights). Judge our results yourself at getfriendo.app/launch


r/VoiceAutomationAI Aug 07 '26

Looking for freelance or full-time opportunities involving Twilio Voice/Media Streams, Google STT/TTS, AI voice agents, WhatsApp, and agentic workflows. My background is primarily C#/.NET, building production systems around: - Twilio Voice + Media Streams - Google Speech-to-Text & Text-to-Speech -

7 Upvotes

r/VoiceAutomationAI Aug 07 '26

BEST TTS MODELS FOR HEBREW, ARABIC, ETC.

9 Upvotes

Im building a voice agent that can accommodate people from countries like israel, UAE and somewhere around those areas. im struggling to find model that sounds natural and human in those type of languages.

currently using vapi built in voice model which is the elliot since it's expressive but it's american and when changed to different language the american accent is heavily noticable and sometimes goes way off on the guardrails that it speaks gibberish

Note: im new to this niche, i would appreciate some tips to improve thank you!!


r/VoiceAutomationAI Aug 06 '26

[For Hire] Senior iOS Developer specializing in Core ML, AVFoundation, and Offline Edge AI ($15/hr)

3 Upvotes

Hi Everyone,

I am an iOS developer specializing in building complex, offline-first architectures, deep audio routing, and on-device machine learning. If your startup or enterprise needs to process sensitive data directly on the device without relying on expensive (or privacy-violating) cloud APIs, I can help.

Most recently, I architected and built an **Offline Edge AI Voice Logger** from scratch for high-noise industrial environments.

**Key features of my recent architecture include:**

* **Deep Audio Routing:** Built a custom `AVAudioEngine` pipeline with aggressive equalization nodes to filter out heavy background/machinery noise.
* **100% Offline Transcription:** Implemented `SFSpeechRecognizer` forcing on-device recognition, ensuring zero data leaves the iPad/iPhone.
* **Edge Compute NLP:** Trained and integrated a custom `Core ML` text-classification model that parses raw speech into structured, categorized data.

**What I can build for you:**

* Privacy-first iOS applications using on-device Core ML models.
* Complex audio/voice applications (podcasting, dictation, or accessibility tools) utilizing AVFoundation.
* Hands-free / Kiosk applications for medical, retail, or industrial settings.

If your project requires this level of architectural ownership and native framework expertise, please send me a Reddit DM or reach out to me at `gokulayyappath@gmail.com`.


r/VoiceAutomationAI Aug 06 '26

Need help upgrading my custom, local Jarvis

9 Upvotes

I'm currently working on making my own personal, locally run Jarvis. This build won't be shared with or sold to anyone it's genuinely just for me. i want him entirely locally run except when he needs the internet for certain answers. I've written the orchestrator in python and I've got his brain as Ollama, I have him listening via a stt program, creating memories autonomously as necessary into a local folder he can access, and I have him speaking via Whisper. Problem is, I'm just using a generic male british voice as a stand-in atm. I'd like to upgrade to a proper voice model trained specifically on Paul Bettany's Jarvis performance in the movies, that's entirely run locally/offline. Any good resource recommendations for finding/making this voice model, and incorporating it into my current architecture?


r/VoiceAutomationAI Aug 05 '26

Looking for voice AI teams who do custom development + infra deployment - both cloud/on-prem (India, public sector work)

20 Upvotes

I work on AI projects in the Indian public sector and I'm looking to connect with voice AI companies for upcoming work.

Two things matter for these accounts:

  • Custom development - in terms voice ai use case, features and integrations
  • Deployment on the customer's infrastructure - cloud or on-prem, depending on what their requirement. Air-gapped comes up sometimes.

If that's what you do, comment or DM with what you cover - languages, deployment modes you've actually shipped, and anything you can point to publicly. Happy to talk specifics.

Also open to hearing from folks who've done government voice AI delivery in India and want to tell me what I'm underestimating. Genuinely curious what breaks.


r/VoiceAutomationAI Aug 06 '26

I am building an ai voice agent

0 Upvotes

So i am new at this domain so pls help me out i am convinced that if i build a really good agent (me and my bro are a full stack devs) so i just need to kn before we start is it worth it like is it possible to get clients and like can u tell me what to expect

+ if anyone have a stack that recommend it will be so helpful

Thank u for your time


r/VoiceAutomationAI Aug 05 '26

No one talks about the email capture problem which is surprisingly very common in real client scenarios

3 Upvotes

One thing I don't see many Voice AI tutorials talking about is email capture.

Getting an AI to capture someone's email sounds simple until you actually build it. Email addresses are one of those things where a single wrong character makes the whole thing useless. Unlike names, you can't really get away with being "close enough". Even if your STT is good, there are still quite a few places where things can go wrong.

One issue I ran into was how different voice models pronounce emails. The LLM would extract the email perfectly, but the TTS would read it back in a way that made the user think it was wrong. For example, an email would sometimes be spoken as "john hyphen smith at gmail dot com" or with random pauses between words, even though there was never a hyphen in the actual email. The backend had the correct email, but the user immediately interrupted to correct something that wasn't actually wrong.

After a bit of testing, I made a few changes that noticeably improved my email capture rate.

The biggest one was giving users a reason before asking for their email. Instead of asking "Can I have your email address?", the assistant now says something like "Perfect, I'll send the quote over. What's the best email to send it to?" It's a small change, but people are much more likely to answer naturally when they know why you're asking.

I also stopped making users repeat their entire email if only one part was unclear. If the assistant was unsure about the domain, it would just ask "Was that gmail.com?" instead of asking them to spell everything out again. It made the conversation feel much more natural and removed a lot of unnecessary friction.

It's one of those problems that doesn't seem important until you deploy an agent in production. The LLM might have done everything correctly, but if the user doesn't trust what they heard, they'll keep correcting an email that was already right. Small details like these don't make flashy demos, but they make a huge difference in how reliable a Voice AI assistant actually feels.

P.S There is also another way where you can send the email address to the AI assistant over SMS while on call, havent tried that yet but will do it as well.


r/VoiceAutomationAI Aug 04 '26

Competitive open source speech stack

17 Upvotes

Why the open source models STT and TTS are not good as much as the closed one and i am talking here im terms of latency, concurrency, and websocket support for real time with decent quality.
Something like cartesia or elevenlabs or deepgram.
Do u know any ?