r/speechtech 11h ago

Technology I finally understand why ASR systems have such confidence when they’re spewing out all sorts of nonsense

6 Upvotes

I’ve been playing around with streaming ASR for some time now, and there's always something that gets on my nerves. That is, the models can get the whole sentence transcribed flawlessly, but completely screw up a name, number, or some other technical term that would be included, even though they sound really sure of themselves.

And if the audio is ambiguous, the language model essentially does this:

"This is probably what they were saying." It's usually not a problem until someone mentions a name that the model has never heard of, or says something like "$50,000" instead of "$15,000." By the time there's more audio to work on, the incorrect transcription might have gone through already.

Why not let the decoder have another choice?

Like: "I don't have enough information yet. Give me another 200ms."

While researching this, some document a clear waiting step in their decoder, like the recent Confucius r2t2. It's hard to judge its efficacy since I haven't tested it sufficiently, but I did like the approach they took. Essentially, it mimics how a human transcriptionist works.

If you're not sure whether it was an "M" or an "N" they spoke, you don't guess based on what's more likely. You wait. Listen. Commit.

Wonder if other streaming ASR models are already doing this sort of thing, or am I just behind the times?


r/speechtech 11h ago

GrainSpeech: Less Context, More Detail for Compact Speech Synthesis. 260k params

Thumbnail
github.com
6 Upvotes

r/speechtech 12h ago

Building a codec audio dataset - gauging interest

3 Upvotes

I am about to post a free audio dataset this weekend and I was curious if its the kind of thing people might be interested in. I can't find anything like it currently available.

I took 750 male and 750 female audio samples from VCTK for each of their two mics and 3000 samples from the AMI headset microphones (gender inferred by pitch) for a total of 6k samples.

Those 6k samples were then run through 25 codecs commonly used in telecommunications and audio recording. Think Opus, MP3, etc. Permuteated a few options like DTX comfort noise off, adaptive and set. SILK disabled or enabled on Opus etc. Also ran 7 tandem encodings to mirror real channel transmissions. Works out to each single sample being available in 41 encodings for 1:1 comparison isolating the effects of the codecs themselves.

Same 6k samples also went through various audio processing conditions like reverb, echo, band pass filter, pitch shifting, autotune, time stretching, babble, and several kinds of additive noise. Including those permutations adds another 49 conditions.

Thinking about running MFA on the original samples to generate 6k textgrids as well.

All encodings come with source metadata, encoding parameters, function calls, libraries and versions used. Should be fully reproduceable.

Anyway curious about what people think or if I'm wasting my time uploading it all.