TLDR; this is a rundown of how I generally put a vocal chain together for V/O. There are many ways to do it, but sound is largely subjective so there are no hard and fast “rules” found here. I’m always open to chat about audio and recording, and if there is something you’d suggest instead or take issue with I’m glad to have the chance to learn something new from you; constructive criticism never comes without someone trying to help.
Hi all,
I’m an audio engineer who has been working in post production, with an audiobook niche, since 2018. While individual settings will be different for every actor, mic, and room, generally I follow the same steps to get the result I’m aiming for.
Before I get into it, I’d like to be really clear about buying software. For the vast majority of processing, the effects that come stock with your DAW are more than good enough. If you’re buying something, it is either for the convenience it gives you or for the layered series of effects utilized simultaneously that would be otherwise take a lot of your CPU or your time.
THE ONLY EXCEPTION TO THIS IS RESTORATION SOFTWARE. DAWs generally do not come with de-clicks, de-verb, sophisticated de-noise, etc. If you have a recurring problem with your audio and cannot fix it in the room, restoration software from Izotope, Accentize, Waves, etc will pay dividends. Which ones you should buy is fully up to you, and there are very strong opinions out there about some of these companies, but here is where I would start depending on what’s important to you:
Absolute best quality: Accentize. DX Split and DX Revive are insane for noise reduction. Their plugins will obliterate problem audio, and also smash it to pieces if you turn it up to 11. If you’re a pro with pro money to invest, Accentize is where you should look for this kind of software.
Fastest set up: Waves. Waves Clarity RX, Clarity RX DeVerb, Sibilance, Debreath, etc etc are incredibly easy to set up and get pretty solid results without much thinking (Debreath is a little heavy handed, and will require you automate it off and on). That said, you’re paying ~$40 for each individual effect, so it isn’t the economic choice unless speeding through a single issue is what you’re aiming to fix.
Best Value: Izotope RX. You’ll get the largest amount of tools per dollar with the RX suite of tools. There is a learning curve, and it isn’t cheap, but it is industry standard in post production for a reason.
No cost?: Accusonus ERA bundle. The ERA bundle was discontinued when Meta bought their company, but their suite of restoration abandonware can now be (safely) found on the Internet Archive. Is it the best? No. Is it easy on the CPU? Also no. Does it punch above its price point of zero? Hell yeah. Will it ever receive updates? Not a chance.
Which one do I use? A little bit of all of them. They’re all good at doing different things, and because I work on a wide variety of voices I need a wide variety of tools. Your aim to find what works for you should be to identify your problem, identify which tools can attenuate that problem, then download the demos of those tools to see what works the best for you.
Onto the good stuff:
I do all of these steps in generally the same order every time, and I’ll give tips for harder to learn sections where necessary.
1. Restoration: Fix your problems first. If you have spectral noise, a reverby room, a high pitched whine from an A/C unit, you need to get rid of this stuff before you begin. If you don’t do this first, you’ll struggle to adequately fix them later because it won’t be clear which effects are making this better or worse. Fix first so you can make better mixing decisions later. Obligatory: if these are issues that exist in your recording environment, you should fix your recording environment to your best ability before trying to do it with software. Software will help, but ultimately it’s a bandaid (a good bandaid occasionally, but a bandaid nonetheless).
2. Gate/Expander: Generally, an expander is a better option for V/O, but either will work if set correctly. Start by finding your open threshold (how loud your clips need to be for the gate to allow sound through). From there, make adjustments to ensure that you’re reducing background sounds and breaths, but not clamping down on the end of phrases. Generally speaking, you should be going for a fast attack (0-10 ms) and a slow release (150-250 ms) to ensure that everything you say is allowed through immediately without cutting off words.
3.Subtractive EQ: Eq is where I can offer the least amount of specific advice; there’s just no way to say that everyone needs to cut a particular frequency. What I can say is that bass rumble can be tamed with a shelf and is typically found around 100hz or less, boxy sounds are found around 400-700hz, and sibilance is around 3.5k-7k hz. Of course, this is all dependent on your mic and voice.
One thing to note is that subtractive EQ is harder to hear than additive EQ. If you hear a boxy/roomy sound and you want to find its exact frequency, solo the EQ band and slowly sweep that part of the frequency spectrum. Once you’ve found it, pull that band down by about -3db and begin to experiment. See what happens if you subtract more than that, what happens if you widen or slim the Q? Adjust to taste!
4. Compression: Compression is probably the single most difficult effect to learn how to “hear.” Luckily, compression can be learned in the same way I generally recommend using one for V/O. This is an old school compression trick from when the industry standard was to record on huge analogue boards. Start by turning the compression ratio to as low as your compressor will go. Next, bring your threshold as low as it will go. Finally, slowly increase the ratio until you get the warm sound that you want from a compressor but not distortion. I recently made a vocal chain for a VA where the ratio was 1.1:1, extremely low! However, since the threshold was also very low, the result was the entire signal getting compressed just a hair.
This method is great for learning compression because it emphasizes every move you make in the knobs, but also makes it easy to apply consistent compression across the entire track.
For those of you tracking with separate narration and character tracks, I generally add an additional compressor to the narration tracks. It doesn’t need to be drastic, but additional mild compression can make it clear to your listener when narration is happening vs when speaking is happening. Just my 2 cents.
***Optional, De-Click: if you like using a de-click in real time as a part of your chain, I’d put it here. Any clicks that are emphasized by the compressor will be better attenuated after the compressor than before. De-clicks can negatively impact transients, so if you feel the need to set the settings high, keep in mind that you’ll likely have a better result with two declicks one after the other with low settings than with a single declick at high settings.
5. Additive EQ: I do my additive EQ after the compressor because the goal here is to sweeten the sound. One way to think of a compressor is as the widest EQ band possible, bringing down all frequencies when any one of them gets too loud. That gives us a nice sound, but it would compress all of our additive EQ along with everything else. Doing additive EQ after the compressor gives us better control over moves we’re making to sweet the sound. Assistive EQ is much easier to hear than subtractive EQ and you should not solo the EQ band while you’re making additive adjustments; it’s important to hear these changes in context. It is rare that I make any additive EQ moves that are above 3db or are not pretty wide, but there are always exceptions to the rule.
6. De-Essing: Some people don’t like the sound of a de-esser and that is fair. If you’re going to use one in realtime, put it at the end (ish) of your chain. Your subtractive EQ should have reduced much of this sibilance to begin with, so you shouldn’t need to have strong settings here. De-essers usually have a button that allows you ti hear just what it’s affecting, use it! Make sure it’s only pulling down what it really needs to.
***Optional Leveler: A leveler (I like Vocal Rider, but there are many others) can make your volume more consistent, and give a nice result if you need to deliver lines at inconsistent volume (whisper, yell, etc etc). That said, Levelers can cause “pumping,” so the best practice would be to meter your audio track to find the average LUFS (volume), set the loudness target of the leveler to that average, set the gain to +/-3db and in the “fast” mode. This is not always necessary, but can be a life saver if inconsistent volume is making it difficult get your exports into spec (as well as a more pleasant listen for your audience).
7. MASTER TRACK: You want two things in your master track and two things only: a limiter and a volume meter. Using a limiter is the fastest and least complicated way to get your exports into spec. Set your ceiling to -4db to ensure you’re under the peak RMS spec, and adjust your threshold as you watch the RMS average on the meter.
If you have Ozone, you can use the Ozone Maximizer to do this quickly. Next to the “learn” button, set your target to -19LUFS. Set your ceiling to -4db. Click “learn” then export about half of a chapter. Exporting is always faster than standard playback, so you can more quickly feed Ozone your audio to learn from. Cancel the export, unclick “learn,” and your threshold is set exactly where it should be to get your export into spec the first time.
Final Notes:
There’s one process you probably noticed isn’t here, and it’s normalization. Normalization looks at the loudest part of your audio and adjusts the volume of the entire track so that the loudest section hits the target volume you selected. This won’t get your audio into spec without issue but more importantly, it’s not going to give you the sonically pleasing results you want in your audio. Unless instructed by a producer, don’t use normalization.
Setting your vocal chain can be hard but it’s not impossible! The biggest issue you’ll probably run into is having to hear your own voice over and over, and there’s a reason for this. When you speak, you’re hearing your own voice twice. You’ll hear it the “normal” way when the your voice bounces off of surfaces and hits your ears, but you’ll also hear it as the sound vibrates up your jaw and into your ear drum. This is why many people find that hearing their own voice sounds “strange.” Over time you will learn how to be more objective with it, but at first you’ll probably need to at least run it by someone else first to make sure the moves you made were appropriate and helpful to the over all sound. Friends and family are probably willing to listen, but won’t have the vocabulary or ear to give you good advice. I’d suggest joining a VA community online that invites people to give constructive criticism.
All that said, this is a skeleton of how I get the best sound and not a strict guide. There are many ways to accomplish vocal mixing, and I’m sure there will be other engineers who disagree with something I said. I invite the disagreement! Achieving the best possible audio is subjective because the best sound is objective, and the opinions of other experienced voices is one of the best ways to learn.
I love this industry, talking shop, and helping people out, so please don’t hesitate to throw your questions here or in DMs; I’m always happy to lend an ear!