audio-branding-and-storytelling
Best Practices for Mixing Podcasts With Multiple Audio Sources
Table of Contents
Mixing a podcast with multiple audio sources requires more than just sliding faders up and down. When you combine multiple microphones, remote guest tracks, music beds, and sound effects, each element competes for space in the frequency spectrum and dynamic range. A well-mixed podcast feels cohesive, maintains consistent volume, and keeps the listener’s attention focused on the content. This guide covers the essential techniques and workflows for mixing multi-source podcasts, from initial gain staging to final export, ensuring professional results every time.
Gain Staging and Signal Flow
Before any creative mixing begins, you must establish clean signal chains for each audio source. Gain staging is the process of setting optimal levels at every stage—from microphone preamplifier to interface input to DAW track fader—so that noise and distortion are minimized while headroom is preserved. Start by setting the preamp gain so that the loudest part of the dialogue peaks around −12 dBFS (or −18 dBFS in 24-bit recording). This leaves ample headroom for dynamic peaks and prevents clipping. Avoid the temptation to “record hot”; modern 24-bit recording allows quiet signals to be raised later without introducing audible noise. Once levels are set in the DAW, your track faders should be adjusted for balance, not for achieving overall loudness. For podcasts with multiple hosts, use busses or groups to control the relative level of all voices together, and apply processing (EQ, compression) to the group for consistency.
Signal Routing Best Practices
- Use separate tracks for each microphone, each remote guest (e.g., via Zoom or SquadCast), and each music or effects source.
- Label and color-code tracks clearly to avoid confusion during editing.
- Set track faders to unity (0 dB) during gain staging, then adjust send levels or bus faders for overall balance.
- Insert a high‑pass filter (80–100 Hz) on all voice tracks early in the chain to remove rumble, mic handling noise, and low‑end room tone.
Room Acoustics and Microphone Placement
The quality of your raw tracks determines how much mixing work you will need. Even the best EQ and compression cannot fix a recording that was made in a reverberant room or with an improperly positioned microphone. For in‑studio recordings, treat the room with absorption panels at early reflection points and behind the speaker. Place microphones six to eight inches from the talent’s mouth, slightly off‑axis to reduce plosives and sibilance. Use pop filters and shock mounts. For remote guests, coach them on microphone placement and encourage them to record in a quiet, small room with soft furnishings. When you receive remote tracks, listen for consistent level and low background noise; if necessary, apply noise reduction (using tools like iZotope RX) before mixing.
Understanding Your Audio Sources
Each audio source has unique tonal and dynamic characteristics. Host voice tracks recorded on a dynamic microphone (e.g., Shure SM7B) often have a warm, focused midrange but lack air. Guest tracks recorded on condenser microphones may have more high‑frequency detail but also more ambient noise. Background music can be dense in the low‑end and midrange, masking dialogue if not managed. Sound effects (applause, transitions) need their own space. Before mixing, listen to each source in solo to understand its natural balance and flaws. This knowledge informs your EQ and compression decisions: for example, you might cut 400 Hz on a boxy‑sounding microphone, or add a gentle high‑shelf boost to a dull guest track.
Setting Proper Levels and Headroom
Initial level balancing is the foundation of a good mix. With all tracks playing, set the faders so that the dialogue sits prominently in the center, around −12 to −9 dBFS on the master bus, with momentary peaks reaching −6 dBFS at most. Background music should be noticeably lower—typically −18 to −24 dBFS or lower—so that it supports the mood without competing. Listen on headphones and speakers at a moderate volume; if you have to strain to hear the dialogue, raise the voices, not the master volume. Use a visual loudness meter (preferably one that measures LUFS) to ensure spoken word averages around −16 to −19 LUFS with a short‑term peak of about −3 dBFS. This leaves room for dynamics and prevents listener fatigue. If you are mixing for broadcast or distribution platforms that require a specific loudness target, set your master limiter accordingly.
Equalization (EQ) for Clarity and Separation
EQ is the most powerful tool for carving space in a multi‑source mix. The goal is to make each element intelligible and balanced without fighting for the same frequencies. For dialogue, start by removing low‑frequency rumble below 80 Hz and apply a gentle cut around 250 Hz to reduce muddiness. Add a subtle presence boost between 2 kHz and 4 kHz for clarity, and a gentle air boost at 8 kHz for brightness—but take care to avoid sibilance. For music beds, use a high‑pass filter at 80 Hz to avoid sub‑sonic buildup, and a notch cut around 250 Hz (the region where voices become boxy) to open up space. If the music competes with dialogue, apply a gentle wide cut in the 1–4 kHz range. Sound effects, such as intros or transitions, can be EQ’d to feel bigger by boosting low end and rolling off highs, while keeping them short to avoid cluttering the mix.
EQ Workflow Tips
- Always cut before boosting: subtractive EQ preserves headroom and sounds more natural.
- Use a narrow band to find resonant frequencies by boosting high, then sweep; once found, cut by 2–4 dB.
- Apply EQ to a bus or group when you need all voices to sound coherent (e.g., all hosts on a “VOX” bus).
- Check EQ adjustments in the context of the full mix—soloing can lead to overprocessing.
Dynamic Processing: Compression, Gating, and De‑essing
Podcasts benefit from moderate compression to smooth out vocal dynamics and keep the dialogue consistent. For voice tracks, start with a 3:1 ratio, attack around 10 ms, and release around 50 ms. Set the threshold so that you achieve 2–6 dB of gain reduction on the loudest phrases. This reduces the distance between quiet and loud sections without sounding overly squashed. Use a compressor on each voice track individually, and then optionally a light compressor on the voice bus for cohesion. Avoid heavy compression; it can make the podcast sound fatiguing and unnatural.
If you have multiple people speaking, a noise gate can prevent cross‑talk and background noise from open microphones. Set the gate threshold just above the noise floor, with fast attack (1–5 ms) and medium release (50–100 ms) to avoid chopping words. A gate is especially useful when each host has a dedicated microphone in the same room.
De‑essing is critical for reducing harsh sibilance (excessive “s” and “sh” sounds). Use a de‑esser or a compressor with side‑chain EQ focused on 5–8 kHz. Apply de‑essing after compression, as compression often brings out sibilance. Attack times around 1 ms and release around 10 ms work well.
Balancing Background Music and Sound Effects
Music and sound effects should augment the narration, not dominate it. The most effective technique for keeping music out of the way of dialogue is side‑chain compression (also called “ducking”). Insert a compressor on the music bus and set its side‑chain input to the voice bus (or to all voice tracks summed). When the host speaks, the compressor reduces the music volume by 3–6 dB, with a fast attack (5–10 ms) and a medium release (200–500 ms) so the music swells back subtly during pauses. This creates a polished, radio‑style effect that maintains energy while preserving clarity. For intros, outros, and transitions, use automatic volume automation to draw in music fades—ramp down before dialogue starts, and ramp up after the last word. Sound effects (applause, stingers) should be treated as punctuation: keep them short, EQ’d to fit, and no louder than the dialogue.
Automation for Transitions and Emphasis
Manual volume automation is your best friend for fine‑tuning a mix. Use automation to:
- Pull down a loud cough or chair squeak that wasn’t edited out.
- Emphasize a key punchline or emotional moment by raising the dialogue level 1–2 dB.
- Smoothly fade music in and out during scene transitions.
- Create a slight volume dip when a second host interrupts to reduce buildup.
Monitoring, Reference Tracks, and Listening Checks
Your mix is only as good as your listening environment. Use high‑quality closed‑back headphones during the initial mix to catch subtle details, then switch to studio monitors or consumer headphones (e.g., AirPods) to test translation. Reference tracks—professionally mixed podcasts or radio shows that you admire—are invaluable. Import a reference track into your DAW, match its loudness to your mix, and switch back and forth to hear differences in balance, EQ, and dynamics. This practice reveals whether your mix is too muddy, too bright, or lacking presence. If possible, listen to your mix in a car, on a phone speaker, and through a laptop to check for problematic resonances or imbalanced levels.
Final Mixdown and Export
After you’ve balanced levels, applied EQ and compression, and automated the mix, it’s time for the final stereo bus processing. Insert a light master compressor or limiter to catch occasional peaks and raise overall loudness by no more than 2–3 dB. Target an integrated loudness of −16 LUFS for most podcast distribution (Spotify, Apple Podcasts often recommend −16 to −19 LUFS). Check your true peak levels—keep them below −1 dBTP to avoid distortion on lossy codecs. Export the final mix as a 44.1 kHz, 16‑bit or 24‑bit WAV file for archiving, then encode to 128–192 kbps MP3 for distribution. Always render at least one version with a few seconds of silence at the start and end to prevent clipping on streaming platforms.
Conclusion
Mixing multiple audio sources for a podcast is a blend of technical discipline and creative judgment. By starting with clean gain staging, treating your recording environment, and applying thoughtful EQ, compression, and automation, you can transform a raw multitrack session into a polished, engaging listening experience. Practice these techniques consistently, and use your ears and reference tracks to guide you. The result will be a podcast that sounds professional and keeps your audience coming back for more.
For further reading, check out the production guides at Sound on Sound, the detailed EQ tutorials at Sweetwater, and the loudness standards overview from iZotope.