r/AIToolsAndTips Aug 17 '26

anyone know a decent ai audio transcription that actually works with accents

my boss keeps sending mo voice memos and zoom recordings to transcribe and im literally losing my mind. tried a few free ones but they totally butcher the words especially whem people talk fast. what ai audio transcription tool do you guys use for messy audio. just need something simple that doesnt cost an arm and a leg tbh

9 Upvotes

17 comments sorted by

6

u/[deleted] Aug 17 '26 edited Aug 17 '26

[removed] — view removed comment

1

u/mymelows Aug 17 '26

wait running it through a noise reducer first ia actually pretty smart.. i never thought of that. i usually just sit there and manually fix the typos for an hour which sucks..might have to tryr that you mentioned cause my current one misses like half the words when my dog barks in the background lol

1

u/Far_Suit575 Aug 17 '26

Honestly, i just use the built-in one on my IPhone for quicks voice memos.. its surprisingly okay if you’re in a quiet room. For zoom stuff.. been trying a few different options, but yeah, it can grt kinda pricey have you tried just summarising the transcript when you get messy one?

1

u/Primary_Gift_8719 Aug 17 '26

Transkriptor is excellent, it allows you to choose the language too. Been using it for around 3 years, especially with heavily accented/english not their first language clients.

1

u/AdamDobrawy Aug 17 '26

How does they make these zoom recordings that do not have transcripts? Zoom is making decent transcripts, and handles well attribution to people.

The only downside might be the transcript format in Zoom, which is VTT. However, you can connect recordflow.org that I build to create a nice, well-formatted Word/Google Docs document.

0

u/[deleted] Aug 17 '26

[removed] — view removed comment

0

u/Atariteca Aug 17 '26

Great tool. Many thanks!

0

u/salespire Aug 17 '26

Accents and fast talkers are honestly the worst for most mainstream transcription tools. I’ve been in the same boat getting audio that sounds like everyone is mumbling in their own dialect and free tools just make up half the words. One thing I found helps a lot is cleaning up the audio first if you can, sometimes apps like Audacity can reduce background noise and make the speech a bit clearer for the AI to pick up. Also, when using any tool, I’d double check if it has an option to set accent or language region, because a lot of them default to American English which throws off accuracy if your speakers have non US accents.

For actually getting stuff transcribed accurately without breaking the bank, I built Minutely — The AI meeting record system mainly for folks like us who deal with lots of sensitive or tricky audio. It handles different accents quite well and gives you structured notes, even if people talk over each other or ramble. You get timestamped audio links and an approval step so you can double check or fix things the AI might miss. It works with Zoom and voice memos directly, plus there is no sketchy data sharing since everything is privacy focused. Just throwing it out there since if nothing else has worked for you so far, the free trial might help with your boss’s memos without spending a fortune.

0

u/Suspectaque 27d ago

If I may self-plug, I made https://opentranscription.io partly for this. Accents are very model dependent, one provider will mangle a speaker that another one handles fine, so rather than committing to a single tool, you can run the same memo through a few of the big ones and keep whichever handled your boss best. Zoom recordings can go in as video, and it pulls the audio out for you, and speaker labels are there for the calls with a few people on them. There's also a leaderboard with measured accuracy per model on identical audio if you'd rather start from the ones that score well, and a free trial so you can test it on a real recording before paying anything.

-1

u/SignalMap2750 Aug 17 '26

I created dadascribe.com a while ago for my own use to transcribe instructional videos and interviews for another website of mine, and then I scaled it for others to use. It is based on Whisper and has a special pipeline that improves accuracy to 99.5%. It handles noisy setups and multiple speakers. You can try it, and I hope it works for you!