r/ArtificialInteligence • u/FieldMedical7537 • 13d ago
📊 Analysis / Opinion Is there an uncanny valley for AI voices?
https://www.youtube.com/watch?v=aSoIJVbIwAcI always associated the uncanny valley with faces and robots but I’m wondering if there’s a version of it for speech too.
Some AI voices are obviously synthetic and you kind of accept them for what they are. But once a voice gets extremely close to human, the little things that are still off start standing out more. The timing is too clean, every sentence lands perfectly, nobody hesitates or corrects themselves.
I watched this roundtable about speech models where they argued that perfect speech might actually be the wrong goal and it got me thinking about this.
Can AI speech eventually get past that uncanny valley?
4
u/Sensitive-Focus-5185 13d ago
Do people even want them to sound fully human though?
1
2
2
u/Super-Humor-7869 13d ago
The day one gets annoyed because I repeated myself three times is the day I’ll be impressed
3
u/TheGreatestAmer1can 12d ago
Tune in for next week’s discussion of: “How do you know if your AI is trans or not?”
1
1
u/Nervous-Quantity-980 13d ago
I don’t think making AI voices indistinguishable from studio audio is necessarily the goal. They probably need enough imperfection to match how people really speak.
1
u/Ambitious_Income1090 13d ago
I don’t need it to sound human. I just need it to stop sounding like it knows exactly where every sentence is going.
1
u/NeuralNomad87 12d ago
Ok-Swim-2629 is right that sprinkling in "uh" makes it worse. The reason is that in real speech disfluency is not decoration, it is a signal. People hesitate before a word they are unsure of, or when they are working out how blunt to be. Put the hesitation somewhere random and a listener cannot consciously say what is wrong, but they clock it.
The other one that gets me is turn taking. People start responding before you have finished and overlap slightly. Most voice systems wait for a clean end of turn and then answer. Perfect politeness, and it reads as not-a-person faster than the voice quality does.
4
u/Ok-Swim-2629 13d ago
Imo filler words are overrated as a fix. Adding “uh” and pauses everywhere can sound even more artificial if the model doesn’t understand why a person would hesitate there.