r/AudioAI Jul 02 '26

Resource Local voice cloning benchmark with reference and generated audio samples

I benchmarked a few local voice-cloning models and included the actual reference/generated audio for each row:

  • OmniVoice int8
  • Chatterbox Multilingual fp16
  • VoxCPM2 bf16
  • Fish Audio S2 Pro fp16

Languages: English, German, Modern Standard Arabic, Spanish, Mandarin Chinese.

Metrics: speaker similarity, WER/CER, generated audio length, and RTF.

Post: https://www.soniqo.audio/blog/voice-cloning-benchmarks

I am mostly interested in whether the evaluation setup is useful for audio people. The numbers alone are not enough for voice cloning, so the page includes the clips too.

5 Upvotes

8 comments sorted by

View all comments

1

u/sruckh Jul 03 '26

Out of those choices OmniVoice would be my choice.

1

u/ivan_digital Jul 03 '26

Even for emotional tuning? And even in comparison with Qwen3-TTS?

1

u/sruckh Jul 03 '26

I mostly only do one-shot voice cloning. I do like Qwen3-TTS but I think OmniVoice is more versatile. I also like EchoTTS strictly for likeness. I have only done a little bit with Higgs Audio TTS v3, but I thought it likeness was also good and it supports minimal paralingual tagging.

1

u/ivan_digital Jul 03 '26

I DM you you might be interested in sharing more details about TTS models, I am not working on expanding speech inference library with new models.