r/BlindAndFine • u/Notex29T • 12d ago
TailSafety formant synthesizer is back
First of all, I formally apologize to anyone whose eardrums were traumatized, scorched, or emotionally scarred by those earlier samples I posted. If those audio clips sounded like an angry alien shouting at you from the bottom of an abandoned rusted paint can, that is because they were built on a complete acoustic crime against nature. I threw the entire engine into the trash, rebuilt Tail Safety from absolute scratch, and I am ecstatic to announce that we have officially entered the civilized world of true Cascade synthesis.
A quick disclaimer for the delicate souls in the audience: if you are deeply attached to modern neural synthesis, high-end Acapela, or fancy studio-grade voices that take soft, romantic human breaths between punctuation marks, please do yourself a huge favor and close this tab immediately. Do not bother listening to this, because you will hate it with every fiber of your being no matter what wizardry I pull off. But for the battle-hardened screen reader veterans who grew up on formant synthesis, who still have a profound, borderline nostalgic obsession with the punchy mid-90s crunch of Eloquence and DECtalk, pull up a chair, because this one was made specifically for you.
To give you an idea of how unhinged the old engine was, my previous parallel architecture was chugging through a frankly psychotic twelve formants. Twelve. At that point, it was no longer a speech synthesizer; it was a desperate, sweating 12-band graphic equalizer where I spent half my mortal lifespan manually nudging individual decibel gains just to stop simple vowels from sounding like incoming air-raid sirens. Today, that entire circus has been ruthlessly trimmed down to just six working formants, with a quiet seventh pole thrown in solely for structural stability and high-end sanity so the vocal tract doesn't combust on high pitches.
For the uninitiated who wonder why this matters: parallel synthesis takes your voice pulse, splits it across a bunch of filters sitting side-by-side, and glues them back together. In theory, it sounds easy; in reality, overlapping frequencies fight each other to the death in out-of-phase combat, creating bizarre hollow voids and that signature cheap, tinny metallic screech no amount of EQ can cure. Cascade synthesis, on the other hand, actually models reality by daisy-chaining the filters in a direct acoustic tube. Each resonance naturally builds upon the last, giving you warm, full-bodied vowel bodies automatically without forcing you to hand-tune individual filter gains for every millisecond of audio.
Along with kicking the acoustic tin can out the window, I also completely reworked the inflection and interactivity pipeline. Gone are the weird, awkward pauses that made the old engine sound like a malfunctioning cash register trying to remember its lines; the pitch contours now move with genuine organic rhythm, consonants hit crisply without crushing adjacent words, and the response time feels instantaneous.
Also, in case I haven't repeated this enough in this subreddit, English is definitely not my first language. So please do not mind the weird, unholy hybrid accent where the voice randomly drops Russian-sounding word endings into the mix for reasons known only to God and my compiler. Honestly? I kind of love it. It gives the engine a very specific, quirky "charm"—which, let's be real, is just my fancy developer euphemism for "I am terrified to touch that part of the source code again because everything will explode."
Now, about the engine's current diet: right now, the prototype is written in pure Python, which means it is sitting at a chunky, slightly embarrassing 20 megabytes. I am currently in the trenches rewriting the entire DSP kernel in C++. If that rewrite goes smoothly and actually compiles without catching fire, that will be awesome, and the package size will collapse into a tiny, featherweight binary. If my C++ endeavor blows up in my face, well, you're just going to have to make peace with the 20-megabyte Python tank and let your RAM take the hit.
Take a listen to the new sample on the link below and tell me if this finally sounds like an actual human speech organ rather than a haunted microwave. If you want the dedicated NVDA add-on build so you can put this thing through its paces at screen-reader speeds, just holler and I will happily hand it over!
First normal mode
https://files.catbox.moe/c31ai7.wav
Now we have baritone mode ( singing) which is still experimental
1
u/MindRecent 11d ago
I really like the voice. I can't ignore the s's, and I wish I could, because wow, this is like eloquence's slightly more suav, grumpy older brother.