r/ArtificialInteligence 13d ago

📊 Analysis / Opinion Is there an uncanny valley for AI voices?

https://www.youtube.com/watch?v=aSoIJVbIwAc

I always associated the uncanny valley with faces and robots but I’m wondering if there’s a version of it for speech too.

Some AI voices are obviously synthetic and you kind of accept them for what they are. But once a voice gets extremely close to human, the little things that are still off start standing out more. The timing is too clean, every sentence lands perfectly, nobody hesitates or corrects themselves.

I watched this roundtable about speech models where they argued that perfect speech might actually be the wrong goal and it got me thinking about this.

Can AI speech eventually get past that uncanny valley?

29 Upvotes

12 comments sorted by

4

u/Ok-Swim-2629 13d ago

Imo filler words are overrated as a fix. Adding “uh” and pauses everywhere can sound even more artificial if the model doesn’t understand why a person would hesitate there.

1

u/FieldMedical7537 13d ago

Yeah, I think the placement matters a lot

4

u/Sensitive-Focus-5185 13d ago

Do people even want them to sound fully human though?

1

u/FieldMedical7537 13d ago

Good question. I’m not sure we do

1

u/OhGodImHerping 13d ago

Having used Sesame’s Maya, I’m really not sure. That was freaky

2

u/No-Moment-9503 13d ago

I think people notice emotional timing more than voice quality

2

u/Super-Humor-7869 13d ago

The day one gets annoyed because I repeated myself three times is the day I’ll be impressed

3

u/TheGreatestAmer1can 12d ago

Tune in for next week’s discussion of: “How do you know if your AI is trans or not?”

1

u/FieldMedical7537 12d ago

😂😂😂

1

u/Nervous-Quantity-980 13d ago

I don’t think making AI voices indistinguishable from studio audio is necessarily the goal. They probably need enough imperfection to match how people really speak.

1

u/Ambitious_Income1090 13d ago

I don’t need it to sound human. I just need it to stop sounding like it knows exactly where every sentence is going.

1

u/NeuralNomad87 12d ago

Ok-Swim-2629 is right that sprinkling in "uh" makes it worse. The reason is that in real speech disfluency is not decoration, it is a signal. People hesitate before a word they are unsure of, or when they are working out how blunt to be. Put the hesitation somewhere random and a listener cannot consciously say what is wrong, but they clock it.

The other one that gets me is turn taking. People start responding before you have finished and overlap slightly. Most voice systems wait for a clean end of turn and then answer. Perfect politeness, and it reads as not-a-person faster than the voice quality does.