audio-branding-and-storytelling
How to Use Compression to Achieve Clearer, More Consistent Audio for Acx
Table of Contents
Why Compression Is Non-Negotiable for ACX Audiobook Success
Recording audiobooks for ACX (Audiobook Creation Exchange) demands a level of audio polish that separates professional narrators from amateurs. The platform's strict technical requirements around noise floor, peak levels, and consistent loudness mean that raw recordings rarely pass muster without processing. Among the essential tools in your chain, compression stands out as the most impactful for turning an uneven, dynamic performance into a smooth, listenable track. When applied correctly, compression compensates for the natural variations in your voice—the shifts in emphasis, the drop in volume at the end of a sentence, or the subtle changes in proximity to the microphone—and delivers a consistent signal that keeps listeners engaged for hours.
ACX listeners expect a uniform listening experience. A chapter that fluctuates between hushed whispers and booming emphasis forces constant volume adjustments and breaks immersion. Compression solves this by narrowing the gap between the quietest and loudest parts of your recording. This does not mean crushing the life out of your performance; rather, it means applying controlled, transparent gain reduction that feels natural. Mastering compression is a skill that directly translates to fewer retakes, faster editing, and a higher likelihood of passing ACX quality control on the first submission.
How Compression Shapes Your Audiobook Recording
The Core Mechanism: Dynamic Range Control
Dynamic range refers to the difference between the softest and loudest sounds in your recording. A typical speaking voice can have a dynamic range of 20 to 30 decibels during normal narration, and this can expand significantly during emotional passages. ACX requires that your final audio stay within a specific loudness window—generally -18 dB to -23 dB RMS or LUFS with peaks no higher than -3 dB. Without compression, you either risk clipping on loud peaks or force your quiet passages to dip below an audible level when normalized.
Compression works by automatically reducing the gain of signals that exceed a set threshold. When the input signal goes above this threshold, the compressor attenuates it by the amount defined by the ratio. For voice work, a ratio of 3:1 means that every 3 dB above the threshold becomes only 1 dB of output. This creates a tighter, more controlled waveform that is far easier to level out during post-production. The result is a recording that requires less manual volume automation and delivers a more professional sound out of the gate.
Key Parameters Every Narrator Must Understand
Compression is not a one-knob solution. Getting usable results requires understanding four interrelated controls: threshold, ratio, attack, and release. Each parameter shapes the compressor’s behavior and, by extension, how your voice will sound in the final mix.
- Threshold: The level at which the compressor starts working. Set this too low and everything gets squashed, creating a lifeless, fatiguing sound. Set it too high and the compressor never engages during normal speech. A good starting point is to set the threshold so that the compressor activates only on the louder syllables and phrases, typically around -18 dB to -24 dB for voiced narration.
- Ratio: Controls how much gain reduction is applied once the signal crosses the threshold. A ratio of 2:1 is subtle and preserves natural dynamics. A ratio of 4:1 is more aggressive and is often used for radio-style voice work. For audiobook narration, 2.5:1 to 3.5:1 is a safe range that provides consistency without sounding processed.
- Attack: How quickly the compressor responds once the signal exceeds the threshold. Fast attack times (1–5 milliseconds) catch transients like hard consonants (P, T, K) and can make the recording sound smooth, but too fast can dull the natural punch of your voice. Slower attack times (10–30 milliseconds) allow the initial burst of a sound to pass through before compression kicks in, preserving clarity and articulation.
- Release: How long the compressor takes to stop applying gain reduction after the signal drops below the threshold. For voice, a release time of 40–80 milliseconds works well for natural speech rhythms. Too short (under 20 ms) can cause pumping and breathing artifacts. Too long (over 150 ms) can make the compressor slow to respond to dynamic shifts, leading to inconsistent levels.
Getting these parameters to work together is the art of compression. Start with a moderate ratio (3:1), a threshold that catches the peaks of normal speech, a medium attack (10 ms), and a medium release (60 ms). Adjust one parameter at a time while listening critically. Over time, you will develop an intuitive feel for how each control shapes your voice.
Building a Compression Workflow for ACX Standards
Step 1: Record with Good Gain Staging
Compression cannot fix a poorly recorded signal. Before you even open a compressor plugin, ensure your recording level is healthy but conservative. Aim for an average level around -18 dBFS on your digital audio workstation (DAW)’s meter, with peaks no higher than -6 dBFS. This headroom gives the compressor room to work without introducing distortion or hitting the ceiling. If your input level is too hot, the compressor will constantly work to control peaks, leading to a squashed and unnatural sound. If it is too low, you will need excessive makeup gain, which amplifies noise and can introduce hiss.
Step 2: Insert Compression Early in the Signal Chain
Place your compressor after any noise gate or high-pass filter but before equalization and limiting. This ordering allows the compressor to see a clean signal that has already been stripped of low-frequency rumble and background noise. Placing compression before EQ means the compressor responds to the natural dynamics of your voice rather than being influenced by boosted frequencies. Once compression has evened out the level, you can apply EQ to shape the tonal balance without the compressor overreacting to those adjustments.
Step 3: Set Compression in Context
Never set compression settings while soloed on a single phrase. Play back a full paragraph or a full chapter section. Listen for sections where your voice naturally gets louder—such as dialogue, emotional passages, or direct quotes—and notice whether the compressor is leveling those sections smoothly or causing audible pumping. If you hear the background noise breathing in and out with the compressor’s action, your release time is too fast. If the loud parts still feel significantly louder than the quiet parts, your ratio needs to increase or your threshold needs to be lowered.
Step 4: Use Makeup Gain to Match Loudness
Because compression reduces the overall level of the signal, you need makeup gain to bring the volume back up to a usable level. Set the output gain so that the compressed signal has a similar perceived loudness to the original, uncompressed signal. A useful technique is to bypass the compressor, listen to the original level, then engage it and adjust the output gain until the two sound equally loud. This ensures you are hearing the effect of compression, not a difference in volume.
Step 5: A/B Test and Iterate
Toggle the compressor on and off while listening to a representative section of your recording. The compressed version should sound more consistent, with less variation between soft and loud phrases. The natural inflection and emotion of your performance should remain intact—if the compression is making your voice sound flat, lifeless, or overly processed, dial back the ratio or raise the threshold. Every voice is different, and the “perfect” compressor settings for one narrator may be completely wrong for another.
Advanced Compression Techniques for Audiobook Narrators
Multiband Compression for Problem Frequencies
Standard compression works across the entire frequency spectrum. Multiband compression splits the signal into separate frequency bands (low, mid, high) and compresses each band independently. This is useful for narrators who have specific issues in certain ranges. For example, a voice with excessive bass buildup on certain vowels can be tamed by compressing only the low-mid band (200–500 Hz) without affecting the clarity of the upper frequencies. Similarly, sibilance issues (exaggerated S and SH sounds) can be addressed with a de-esser, which is essentially a multiband compressor focused on the 5–10 kHz range. For ACX narrators recording in less-than-ideal rooms, multiband compression can help mask room resonances that a standard compressor would amplify.
If you choose to use a multiband compressor, use it sparingly. Overuse leads to an unnatural, “telephone” quality that fatigues the listener. A light touch—1 to 2 dB of gain reduction in the problem band—is often enough to clean up the sound without making it sound processed.
Serial Compression: Layering for Transparency
Instead of using one compressor with a high ratio and aggressive settings, try using two compressors in series, each with a lower ratio and gentler settings. This approach, known as serial compression, achieves consistent leveling with less audible processing. For example, a first compressor with a ratio of 2:1 and a low threshold can catch broad dynamic shifts, while a second compressor with a ratio of 3:1 and a higher threshold can tame the remaining peaks. The combined effect is smooth and transparent, preserving the natural quality of your voice while achieving the level consistency that ACX requires.
Serial compression is particularly effective for narrators with highly dynamic performances—those who speak softly during narrative passages and then project strongly during character dialogue. The first compressor handles the large-scale dynamics, and the second compressor polishes the small details without introducing artifacts.
Limiting as a Safety Net
After compression, a limiter can serve as a final safeguard against stray peaks that exceed your target level. Set the limiter’s ceiling to -3 dBFS (the maximum peak allowed by ACX) and use a fast attack time (1 ms or less) to catch any transient that the compressor missed. The limiter should only activate occasionally; if it is constantly engaged, your compressor settings are too conservative and you need to adjust the threshold or ratio. A well-set limiter is invisible to the listener and ensures your file will not fail ACX quality checks due to clipping.
Common Compression Mistakes and How to Fix Them
- Over-compressing for loudness: The most common error is chasing loudness by using a high ratio and low threshold. This produces a flat, fatiguing sound that lacks dynamics and listener engagement. Instead, focus on consistency rather than maximum volume. ACX has specific loudness targets; exceeding them is as problematic as falling short.
- Ignoring the noise floor: Compression reduces dynamic range, which includes the noise in your recording. If you have a noisy room, a humming computer, or a microphone with high self-noise, compression will make that noise more audible between words. Always apply a noise gate before compression if your recording has any background noise, and ensure your recording environment is as quiet as possible.
- Using defaults without listening: Preset compressor settings labeled “vocal” or “voiceover” can be starting points, but they are rarely optimal for your specific voice and recording conditions. Always adjust the threshold and ratio to match your recording level and performance style. A preset designed for a booming radio voice will sound terrible on a close-miked, intimate narration.
- Setting attack and release too fast: Fast attack and release times can create a pumping effect as the compressor rapidly engages and disengages with each syllable. This is especially noticeable on plosive sounds (P, B) and sibilants. Allow the compressor to breathe with your natural speech rhythm by using moderate settings and listening for artifacts.
- Neglecting makeup gain: Compressors reduce level; therefore, you need to compensate. Failing to add sufficient makeup gain will result in a quieter file that requires normalization later, potentially introducing quantization noise or forcing excessive gain that amplifies noise.
Compression and ACX Technical Requirements
ACX has published detailed audio submission requirements that every narrator must meet. These include:
- Peak levels: Must not exceed -3 dBFS.
- Noise floor: Must be at or below -60 dBFS (excluding room tone).
- Loudness: RMS or LUFS levels should average between -18 dB and -23 dB with a maximum of -20 dBFS RMS for consistent volume.
- Uniformity: The volume should be consistent across the entire recording, with no significant level shifts between chapters or within a single chapter.
Compression directly addresses the uniformity requirement. When used in conjunction with a limiter and proper gain staging, compression ensures that your peaks stay within the -3 dBFS ceiling while your average level sits comfortably in the target loudness window. Without compression, you will either have to manually ride the fader for every chapter or risk failing the uniformity check.
It is worth noting that ACX also requires a consistent background noise floor. Compression makes noise issues more apparent, so record in a treated space and use a noise gate set to -60 dBFS or lower to clean up any residual noise between words. For more detailed guidance on meeting ACX technical specs, refer to ACX’s official submission requirements page.
Choosing the Right Compression Tools
DAW-Based Compressors
Every major DAW includes a stock compressor that is capable of professional results. Reaper’s ReaComp is surprisingly flexible for a free bundled plugin and offers multiband compression features. Logic Pro’s Compressor includes multiple circuit models (Platinum, Studio, Vintage) that mimic classic hardware units. Pro Tools users have the stock Dynamics III compressor, which is clean and efficient. Audacity, a popular free DAW among budget-conscious narrators, includes a basic compressor with threshold, ratio, attack, and release controls, though its visual feedback is limited.
If you are just starting out, use your DAW’s built-in compressor until you understand the parameters well enough to know what you want from a third-party plugin. Upgrading too early can mask knowledge gaps and lead to confusion.
Third-Party Compressor Plugins Worth Considering
- Waves RVox: A vocal-specific compressor that simplifies control to three main knobs: Attack, Release, and Compression. Its preset “Narration” mode is a solid starting point for audiobook work.
- iZotope Nectar: An all-in-one vocal processing suite that includes a compressor with visual feedback and preset chains designed for voiceover and narration. Nectar also includes a de-esser and EQ, making it a one-stop solution for ACX processing.
- FabFilter Pro-C 2: A transparent, highly flexible compressor with excellent metering and a “Vocal” style that works well for narration. The sidechain filter is particularly useful for preventing low frequency content from over-activating the compressor.
- Slate Digital FG-Stress: A character compressor that adds subtle warmth and color while compressing. Useful for narrators who want a slightly richer tone without heavy processing.
For a deeper comparison of compression tools and their suitability for spoken-word recording, this review of vocal compressor plugins provides practical insights from experienced engineers.
Adapting Compression to Your Voice Type and Recording Environment
For Deeper Voices
If you have a naturally deep voice with strong low-frequency presence, be careful about compression emphasizing that bass. Use a high-pass filter before the compressor (cutting below 80 Hz) to prevent low-frequency rumble from engaging the compressor unnecessarily. A slower attack time (15–20 ms) will preserve the natural warmth and depth of your voice while still controlling peaks. Avoid ratios above 4:1, as they tend to make deep voices sound muddy and indistinct.
For Higher or Softer Voices
If your voice sits in a higher register or you tend to speak softly, you need more gain reduction to bring up the quiet sections. A lower threshold (around -22 dB) and a moderate ratio (3:1) will raise the perceived level of your narration without making it sound strained. Use a faster attack time (5–8 ms) to catch sibilant peaks that are more common in higher voices. A de-esser in series with your compressor is highly recommended to manage S and T sounds that can become harsh after compression.
For Room-Mic Recording
Some narrators record with a room microphone (large-diaphragm condenser at a distance) rather than a close headset or dynamic mic. This setup captures more room sound and requires careful compression. Use a lower ratio (2:1) and a higher threshold to avoid pumping on room reflections. A multiband compressor can help control the low-mid buildup that is common in untreated room recordings. Always gate the signal before compression to prevent room noise from being amplified.
Practical Workflow Example
Here is a concrete chain that works well for most audiobook narrators recording in a treated home studio with a condenser microphone:
- Recording: Input level averaged at -18 dBFS, peaks at -8 dBFS. No clipping.
- Noise gate: Threshold at -60 dBFS, attack 2 ms, release 100 ms. Catches breath noises and light background hum.
- High-pass filter: Cut at 80 Hz to remove low-end rumble and proximity effect.
- Compressor: Ratio 3:1, threshold -20 dB, attack 10 ms, release 60 ms. Gain reduction reads 3–5 dB during normal speech.
- EQ: Slight shelf boost at 3 kHz (+2 dB) for clarity and presence. Subtle cut at 200 Hz if voice sounds boxy.
- Limiter: Ceiling at -3 dBFS, attack 0.3 ms, release 20 ms. Catches any remaining peaks.
- Normalization: Loudness normalize to -20 dB LUFS using a loudness meter plugin.
This chain consistently produces files that pass ACX checks with room to spare. Adjust the compressor threshold and ratio based on your specific recording level and voice dynamics.
Final Checks Before Export
Before you export your final WAV or MP3 file for ACX submission, run these checks:
- Loudness metering: Use a plugin like YouLean Loudness Meter or the built-in loudness meter in your DAW to verify that the average loudness is between -18 dB and -23 dB LUFS.
- Peak check: Confirm that no peak exceeds -3 dBFS. Zoom in on the waveform to look for any clipped samples.
- Noise floor check: Isolate a section of silence (dead air) and measure the RMS level. It should be at or below -60 dBFS.
- Consistency listen: Play a full chapter at a comfortable listening level. The volume should remain steady from start to finish without jarring shifts during emotional passages.
- ACX ACX Check Tool: Use the free ACX Audio Submission Checker to validate your file before uploading. This tool catches technical issues that your ears might miss.
Moving Beyond the Basics
Once you have mastered single-band compression and are consistently meeting ACX standards, explore dynamic EQ as a complementary tool. Dynamic EQ applies EQ only when a certain frequency exceeds a threshold, which can be used to tame harshness without dulling the entire recording. Similarly, parallel compression (blending a heavily compressed signal with the dry signal) can add thickness and consistency without sacrificing the natural dynamics of your voice. These advanced techniques are best learned gradually, as they require a solid foundation in the core principles covered here.
For further reading on the science of compression and how it interacts with the human ear, Sound on Sound’s comprehensive guide to compression is an authoritative resource that breaks down the physics and practical application in detail. Understanding the theory behind the knobs will make you a more confident and effective engineer for your own voice.
Compression is not a magic fix, but it is the single most powerful tool you have for turning a good recording into a great one. By respecting the parameters, listening critically, and adapting your settings to your unique voice, you will produce audiobooks that keep listeners turned on and ACX auditors satisfied. Practice with intention, and the results will speak for themselves.