audio-tutorials
How to Use Multiband Compression in Audiobook Mastering
Table of Contents
Understanding Multiband Compression in Audiobook Mastering
Multiband compression stands as one of the most precise tools available for audio mastering, particularly when working with spoken word content like audiobooks. Unlike standard single-band compression, which applies gain reduction evenly across the entire frequency spectrum, multiband compression divides your audio into separate frequency ranges. This allows you to compress the low end differently from the mids and highs. Think of it as having a dedicated compressor for the bass, another for the vocal body, and a third for the sibilance region—all working in parallel. This surgical approach gives you control over specific tonal issues without flattening the life out of the narration. For audiobook producers aiming to meet strict loudness standards while preserving clarity and naturalness, mastering this technique is essential.
The Frequency Spectrum of Spoken Word
Before diving into settings, it helps to understand how the human voice occupies the frequency spectrum and which areas cause the most trouble in audiobook production.
- Sub-bass and bass (20 Hz–150 Hz): This region contains room rumble, HVAC noise, and physical thumps from the recording booth. While the human voice has little energy here, untreated low-frequency noise can cause the limiter to pump and create an unstable listening experience.
- Low mids (150 Hz–500 Hz): The fundamental frequencies of most male and some female voices live here. Too much energy in this range creates a boomy or muddy quality. Too little makes the voice sound thin and hollow.
- Upper mids (500 Hz–3 kHz): This is the intelligibility zone. The ear is most sensitive to this range, and consonants like T, K, and P rely on it. Overcompression here can make speech sound aggressive or harsh.
- High frequencies (3 kHz–10 kHz): Sibilance (S and Sh sounds), vocal fricatives, and air live here. Too much can cause listening fatigue; too little makes the narration dull and lifeless.
- Air band (10 kHz–20 kHz): This region adds openness and sheen. For audiobooks, it's often rolled off gently to reduce noise and avoid amplifying microphone hiss.
Multiband compression lets you address each of these zones independently, applying as much or as little gain reduction as needed without affecting the rest of the signal.
Why Multiband Compression Is Essential for Audiobooks
Standard compression has been a staple of vocal processing for decades, but it comes with a fundamental limitation: it reacts to the loudest part of the spectrum. If a narrator's voice has a boomy low end and a sharp sibilant peak, a single compressor will trigger on whichever element is loudest at a given moment. This can lead to pumping where the low end modulates the whole mix, or inconsistent de-essing as the compressor clamps down on every high-frequency transient. Multiband compression solves this by giving each frequency region its own threshold, ratio, attack, and release. Here are the specific benefits for audiobook mastering.
Enhancing Vocal Intelligibility
Intelligibility in spoken word hinges on the clarity of consonants and the balance between the fundamental and formant frequencies. When a recording has uneven frequency response due to microphone proximity effect or room acoustics, certain words can become indistinct. Multiband compression allows you to gently reduce gain in problem areas—like a 200 Hz bump from a close-miked male voice—while leaving the upper mids untouched. This preserves the natural articulation of the narrator without making the processing audible. For example, a 2:1 ratio on the low-mid band with a slow attack can even out chest resonance without flattening the dynamic expression that makes a performance engaging.
Controlling Sibilance Without Dulling the Track
Sibilance is one of the most common complaints in audiobook production. Standard de-essers often work by sidechaining a compressor to a specific frequency band, but they can leave the rest of the high end sounding dull if they are too aggressive. A multiband compressor set to the 4 kHz–8 kHz range gives you fine-grained control. You can set a threshold that only activates during the sibilant peaks, with a fast attack (1–3 ms) and a medium release (50–100 ms) to clamp down on the harshness while allowing the natural air and detail to pass through unchanged. This preserves the narrator's vocal character while eliminating the piercing quality that causes listener fatigue.
Managing Low-Frequency Noise and Rumble
Audiobooks are often recorded in home studios or portable booths where low-frequency noise from traffic, HVAC systems, or nearby equipment can seep into the signal. High-pass filtering removes a lot of this rumble, but sometimes the noise sits in the same range as the lower harmonics of a deep voice. Multiband compression offers a gentler solution: a low-band compressor with a threshold set just above the noise floor will grab the rumble when it appears without cutting into the voice. This can be especially useful for recordings where the noise is intermittent, such as a passing truck or a bump in the room. The compressor acts only when the low-frequency energy exceeds the threshold, leaving the rest of the performance clean.
Ensuring Consistent Loudness Across Chapters
One of the greatest challenges in audiobook mastering is maintaining a consistent loudness level across an entire project, which may have been recorded over multiple sessions with different microphone positions, preamp settings, or even different rooms. Multiband compression helps smooth out these variations by reacting to each frequency band independently. A chapter recorded slightly closer to the mic might have boosted bass; the low-band compressor can rein that in without affecting the mids or highs. Meanwhile, a quieter chapter with a thinner tone can have its upper mids gently enhanced. By automating the threshold or ratio across bands, you can achieve a seamless listening experience that meets loudness standards such as ACX's -23 dB LUFS target without sounding over-processed.
Step-by-Step Guide to Applying Multiband Compression
Now that you understand the "why," let's walk through the "how." The following steps assume you have a multiband compressor plugin inserted on your master bus or a dedicated audiobook mastering chain. Always make adjustments in the context of your full signal chain, including EQ, limiting, and any noise reduction.
1. Choose the Right Plugin
Your choice of plugin will significantly affect your workflow and results. Look for a multiband compressor that offers flexible crossover controls, solo per band, and visual metering so you can see what each band is doing. Here are three widely used options that work well for audiobook mastering:
- iZotope Ozone Dynamics: Part of the Ozone mastering suite, it offers four bands with intuitive crossover control and includes adaptive release modes that work well with speech. The built-in Tonal Balance Control can help you match your audiobook to a target curve. iZotope provides extensive guides on vocal processing that are worth studying.
- FabFilter Pro-MB: Known for its transparent sound and clean interface, Pro-MB allows up to six bands with fully adjustable crossover slopes. The "Brickwall" crossover mode minimizes phase distortion, which is critical for preserving the natural phase coherence of spoken word. FabFilter's video tutorials are excellent for learning advanced techniques.
- Waves C6: A more affordable option that still offers six bands with decent flexibility. It does not have the same transparency as FabFilter or the integrated metering of iZotope, but it can be effective for managing problem frequencies once you are comfortable with the tool.
For beginners, starting with a plugin that includes presets designed for voice or spoken word can give you a foundation to tweak. Just remember to bypass the preset and start from scratch once you understand what each band is doing.
2. Set Your Frequency Bands and Crossovers
For audiobook mastering, you typically need three to four bands. Using more bands increases the risk of phase issues and can make the sound feel disconnected or "multi-band-y." A good starting point for a clean narration is:
- Band 1: 20 Hz–150 Hz (low rumble and bass)
- Band 2: 150 Hz–2.5 kHz (body and fundamental frequencies)
- Band 3: 2.5 kHz–8 kHz (presence and sibilance)
- Band 4: 8 kHz–20 kHz (air and high-frequency noise)
Adjust the crossover frequencies based on your narrator's voice. A deep male voice might benefit from a lower crossover between bands 1 and 2 (around 120 Hz), while a female voice with a naturally brighter tone might need the band 2–3 crossover pushed closer to 3 kHz. Always use linear-phase crossovers if your plugin offers them, as this minimizes phase shift around the crossover points. If you hear a slight "hollow" sound or a dip in frequency response at the crossover, try overlapping the bands slightly or adjusting the crossover slope.
3. Set Thresholds and Ratios for Each Band
With your bands defined, set each band to a moderate ratio between 2:1 and 3:1. This gives you gentle control without aggressive pumping. Then lower the threshold of each band until the gain reduction meter shows approximately 2–4 dB of compression on the loudest peaks. The key here is to listen—not just look at the meters. If you see 4 dB of compression but the narrator sounds squashed, reduce the ratio. If the compressor is barely reacting and you still hear unevenness, lower the threshold.
Band 1 (Low rumble): Use a medium attack (10–20 ms) to let the initial thump of a plosive pass through, and a medium release (100–200 ms) to avoid pumping. Aim for 2–3 dB of gain reduction on the biggest low-frequency peaks. This should clean up rumble without making the bass sound choppy.
Band 2 (Body and mids): This band carries the core of the voice. Use a slower attack (10–30 ms) to preserve the natural transient of consonants, and a release that matches the speech rhythm (50–100 ms for most narrators). You only need 1–3 dB of compression here—too much will make the voice sound thick and static.
Band 3 (Presence and sibilance): This is where you control harshness. Use a fast attack (1–5 ms) to catch sibilant peaks quickly, and a medium-fast release (30–50 ms) so the compressor recovers before the next word. Set the threshold so that only the harshest sibilants trigger gain reduction (around 3–6 dB on the peaks). If you find yourself compressing more than 6 dB regularly, consider adding a dedicated de-esser before the multiband compressor.
Band 4 (Air and noise): This band often requires the lightest touch. Use a slow attack (20–40 ms) to avoid cutting into the natural breath sounds that give life to the narration, and a fast release (20–40 ms). You might only see 1–2 dB of gain reduction, just enough to smooth out any harsh high-frequency content or microphone hiss that survives the EQ stage.
4. Fine-Tune Attack and Release Times
The attack and release times are the most critical parameters for natural-sounding multiband compression on speech. If the attack is too fast, you will squash the transient of consonants, making the narrator sound lispy or distant. If the attack is too slow, the compressor never catches the problem. A good starting point for speech is to set the attack of each band slightly slower than you think you need, then gradually speed it up until you hear the compressor grab the peak without destroying the feel of the word. For release times, aim for the compressor to recover fully between words but not between syllables. Listen to a phrase like "the quick brown fox" and watch the gain reduction meter. You want it to release fully during the pause before "fox" but not during the quick "ck" sound.
5. Verify With A/B Testing and Metering
Always compare your processed signal against the original. Most multiband compressor plugins have a bypass button that lets you toggle the entire effect. A good mastering job should sound clearer, more consistent, and less fatiguing than the raw recording. If you hear pumping, unnatural changes in tone, or a loss of dynamic expression, your settings are too aggressive. Go back and reduce the ratio or raise the threshold on the offending band. Additionally, use a loudness meter to verify that your integrated loudness stays within the target range (typically -23 dB LUFS for ACX or -18 dB LUFS for other platforms). Multiband compression can change the perceived loudness of a track, so you may need to adjust your final limiter gain to compensate.
Common Mistakes and How to Avoid Them
Even experienced engineers can fall into traps when using multiband compression on spoken word. Here are the most common pitfalls and how to sidestep them.
Over-Compression and the "Sausage" Effect
The most frequent mistake is applying too much gain reduction across all bands. When every band is compressed 6–10 dB, the voice loses its natural dynamic contour. The narrator's quiet moments become as loud as their passionate outbursts, robbing the performance of emotional nuance. This is often called the "sausage" effect because the waveform looks like a solid block. To avoid this, limit your compression to 2–4 dB per band as a rule of thumb, and use higher ratios only for specific problem frequencies that appear intermittently.
Wrong Crossover Points
Setting crossover frequencies without listening can create audible artifacts. If your crossover point lands exactly on the formant of a vowel sound—say around 800 Hz for an "ah" sound—you might hear a subtle notch filter effect as the two bands compete. To find clean crossover points, use the solo function on each band and listen for any frequency that sounds emphasized or hollow when you move the crossover. Then adjust in 20–50 Hz increments until the transition is smooth. A good practice is to set crossover points at frequencies where the voice has relatively little energy, such as 150 Hz, 2.5 kHz, and 8 kHz, as mentioned earlier.
Neglecting Phase Coherence
All multiband compressors introduce phase shifts at the crossover points. While these are often inaudible on complex music material, spoken word with its clean harmonic structure can reveal phase issues as a slight "smearing" of the transient or a hollow quality in the voice. Use linear-phase mode whenever possible. If your plugin does not have linear-phase crossovers, try to use as few bands as possible and keep the crossover slopes gentle (12 dB/octave instead of 24 dB/octave) to reduce phase rotation.
Integrating Multiband Compression With Other Mastering Tools
Multiband compression should not work in isolation. For professional audiobook mastering, it fits into a signal chain that typically includes the following stages in order:
- Noise reduction: Remove clicks, pops, mouth sounds, and background noise using a spectral editor or noise gate.
- EQ: Shape the overall tonal balance. Cut low-end rumble, reduce boxiness around 200–400 Hz, and add a gentle shelf boost around 3–5 kHz for presence if needed.
- Multiband compression: Even out the dynamics across frequency bands as described in this guide.
- De-essing: If sibilance is still prominent after multiband compression, use a dedicated de-esser in the 5–8 kHz range as a final cleanup.
- Limiting: Raise the overall level to your target loudness using a transparent brickwall limiter with a ceiling of -1 dB or -0.3 dB to prevent intersample peaks.
- Loudness normalization: Use a loudness meter to verify integrated LUFS and true peak values, adjusting the limiter gain as needed to hit your target.
For more detailed guidance on building a complete mastering chain for spoken word, the ACX (Audiobook Creation Exchange) submission guidelines provide a solid baseline for loudness and technical specs. Additionally, articles from production sound experts can help you understand how these tools interact in practice.
Conclusion
Multiband compression is a precision instrument in the audiobook mastering workflow. When used correctly, it gives you the power to clean up low-frequency noise, control harsh sibilance, enhance vocal clarity, and maintain consistent loudness across an entire project—all while preserving the natural dynamics that make a narrator's performance compelling. Start with the band settings and parameters outlined here, trust your ears more than the meters, and always compare against the unprocessed signal. Over time, you will develop an intuitive sense for how much compression each band needs, and you will be able to handle any recording challenge with confidence. The best mastering is invisible: the listener should never think about the tools, only be absorbed in the story.