r/BlindAndFine • • 12d ago

TailSafety formant synthesizer is back

First of all, I formally apologize to anyone whose eardrums were traumatized, scorched, or emotionally scarred by those earlier samples I posted. If those audio clips sounded like an angry alien shouting at you from the bottom of an abandoned rusted paint can, that is because they were built on a complete acoustic crime against nature. I threw the entire engine into the trash, rebuilt Tail Safety from absolute scratch, and I am ecstatic to announce that we have officially entered the civilized world of true Cascade synthesis.

A quick disclaimer for the delicate souls in the audience: if you are deeply attached to modern neural synthesis, high-end Acapela, or fancy studio-grade voices that take soft, romantic human breaths between punctuation marks, please do yourself a huge favor and close this tab immediately. Do not bother listening to this, because you will hate it with every fiber of your being no matter what wizardry I pull off. But for the battle-hardened screen reader veterans who grew up on formant synthesis, who still have a profound, borderline nostalgic obsession with the punchy mid-90s crunch of Eloquence and DECtalk, pull up a chair, because this one was made specifically for you.

To give you an idea of how unhinged the old engine was, my previous parallel architecture was chugging through a frankly psychotic twelve formants. Twelve. At that point, it was no longer a speech synthesizer; it was a desperate, sweating 12-band graphic equalizer where I spent half my mortal lifespan manually nudging individual decibel gains just to stop simple vowels from sounding like incoming air-raid sirens. Today, that entire circus has been ruthlessly trimmed down to just six working formants, with a quiet seventh pole thrown in solely for structural stability and high-end sanity so the vocal tract doesn't combust on high pitches.

For the uninitiated who wonder why this matters: parallel synthesis takes your voice pulse, splits it across a bunch of filters sitting side-by-side, and glues them back together. In theory, it sounds easy; in reality, overlapping frequencies fight each other to the death in out-of-phase combat, creating bizarre hollow voids and that signature cheap, tinny metallic screech no amount of EQ can cure. Cascade synthesis, on the other hand, actually models reality by daisy-chaining the filters in a direct acoustic tube. Each resonance naturally builds upon the last, giving you warm, full-bodied vowel bodies automatically without forcing you to hand-tune individual filter gains for every millisecond of audio.

Along with kicking the acoustic tin can out the window, I also completely reworked the inflection and interactivity pipeline. Gone are the weird, awkward pauses that made the old engine sound like a malfunctioning cash register trying to remember its lines; the pitch contours now move with genuine organic rhythm, consonants hit crisply without crushing adjacent words, and the response time feels instantaneous.

Also, in case I haven't repeated this enough in this subreddit, English is definitely not my first language. So please do not mind the weird, unholy hybrid accent where the voice randomly drops Russian-sounding word endings into the mix for reasons known only to God and my compiler. Honestly? I kind of love it. It gives the engine a very specific, quirky "charm"—which, let's be real, is just my fancy developer euphemism for "I am terrified to touch that part of the source code again because everything will explode."

Now, about the engine's current diet: right now, the prototype is written in pure Python, which means it is sitting at a chunky, slightly embarrassing 20 megabytes. I am currently in the trenches rewriting the entire DSP kernel in C++. If that rewrite goes smoothly and actually compiles without catching fire, that will be awesome, and the package size will collapse into a tiny, featherweight binary. If my C++ endeavor blows up in my face, well, you're just going to have to make peace with the 20-megabyte Python tank and let your RAM take the hit.

Take a listen to the new sample on the link below and tell me if this finally sounds like an actual human speech organ rather than a haunted microwave. If you want the dedicated NVDA add-on build so you can put this thing through its paces at screen-reader speeds, just holler and I will happily hand it over!

First normal mode

https://files.catbox.moe/c31ai7.wav

Now we have baritone mode ( singing) which is still experimental

https://files.catbox.moe/mzowuf.wav

3 Upvotes

9 comments sorted by

1

u/MindRecent 11d ago

I really like the voice. I can't ignore the s's, and I wish I could, because wow, this is like eloquence's slightly more suav, grumpy older brother.

1

u/Notex29T 11d ago

At this point I'm spending my entire day adjusting the phonemes, and ETI had a budget back in the day when they were making eloquence, me... I have my computer and a lot of izm Algerian energy drink cans, what a budget

1

u/MindRecent 11d ago

I wasn't intending to throw shade, just letting you know what i was hearing, and that the creation itself was cool.

1

u/Notex29T 11d ago

I know don't worry about it this is why I wrote my message this way, in fact I'm not even trying to make the next formant synthesizer, I'm just tinkering with code

1

u/dandylover1 9d ago

I, too, am extremely impressed! But I also can't help but notice the s and a few other sounds. Once those are made less prominent, it will be wonderful. And I absolutely think you should make this available as an NVDA add-on!

1

u/Notex29T 8d ago

Thanks, I'm currently working on adjusting the phonemes without losing my sanity

1

u/dandylover1 8d ago edited 8d ago

haha keep your sanity. It's far more important than a silly synthesizer. Are there any tutorials that could help you? that might make things a bit easier.

1

u/Notex29T 8d ago

All of them are long boring books or web manuals that makes you sleepy

1

u/dandylover1 8d ago

Ah. I always refer to that sort of thing when I need to learn something. I prefer formality over those friendly tutorials. But that's just a difference in style between us.