field-recording-and-soundscapes
How to Use Mid/side Processing to Enhance Dialogue Spatiality
Table of Contents
Understanding Mid/Side Processing
Mid/side (M/S) processing is a technique that breaks a stereo signal into two components: the mid (sum of left and right, also called the mono signal) and the side (difference between left and right). By manipulating these components separately, you can adjust the perceived width, depth, and focus of audio without altering the original stereo image in a destructive way. For dialogue, this tool is invaluable—it lets you center the voice, tighten its presence, and shape the surrounding environment with precision.
The Sum and Difference Matrix
The math behind M/S is straightforward: Mid = L + R, Side = L – R. In practical plugin terms, an M/S encoder applies this matrix, converting the input stereo into two output channels labeled Mid and Side. To return to normal stereo, a decoder reverses the process: L = (Mid + Side) / 2, R = (Mid – Side) / 2. Most M/S plugins handle both encoding and decoding in one interface, but you can also build the chain manually using routing — for example, in a DAW like Logic Pro, Pro Tools, or Cubase.
Real-World Analogy
Think of the mid channel as the voice standing directly in front of you in a room, while the side channel captures the reflections off the walls. If you boost the mid, that central point becomes more dominant. If you boost the side, the room’s acoustics sound wider, more spacious, or even more diffuse. For dialogue clarity, the mid is your anchor; the side gives you control over the spatial context.
Why Dialogue Benefits from Mid/Side Processing
In film, television, game audio, and podcasting, dialogue must remain clear and intelligible even against background music, sound effects, or layering. Traditional stereo EQ and compression treat the entire image equally, which can muddle the voice or create unbalanced width. Mid/side processing addresses these challenges directly.
Centering the Voice
The mid channel carries all the information that is identical in both left and right — which is where the dialogue signal typically lives if it was recorded with a mono source or panned centrally. By applying compression, EQ, or even de-essing only to the mid channel, you can tighten the vocal presence without affecting the stereo effects or ambient width. This results in a voice that cuts through the mix without feeling artificially isolated.
Controlling Spatial Ambience
The side channel contains the non‑identical information — room tone, stereo reverb tails, and any stereo‑placed effects. You can adjust the side level to make the background seem larger or smaller. For a natural, immersive dialogue scene, you might boost the side slightly to give the impression of a real space, but you can also reduce it for a dry, intimate, or direct sound — common in voice‑over or ADR.
Compatibility with Mono Playback
One of the strongest arguments for M/S processing is its mono compatibility. Because the mid channel contains all the mono information, any changes you make to the side (like adding stereo reverb) or to the mid (like compression) will not collapse into phase cancellation when the signal is summed to mono. The dialogue remains centered and intelligible, even on single‑speaker playback systems. This is critical for broadcast, podcasts, and public venues.
Practical Setup in Your DAW
Choosing a Plugin
Many professional M/S plugins exist. Here are reliable options:
- Voxengo MSED – free and simple, does only encoding/decoding (no processing). Ideal for combining with your own EQs and compressors.
- Brainworx bx_digital V3 – includes dedicated mid/side EQs, compressors, and a stereo widener.
- Waves S1 Imager – not strictly M/S but can adjust stereo width; combine with other tools.
- FabFilter Pro‑Q 3 – allows mid/side EQ processing directly in the plugin (no separate encoder needed).
- iZotope RX – offers M/S mode in its EQ and other modules, particularly useful for dialogue and speech.
Download Voxengo MSED (free) – a good starting point.
Step-by-Step Workflow
Follow this process to apply M/S processing to a stereo dialogue track. The steps assume you are using a dedicated M/S encoder plugin followed by separate processors on the mid and side channels – but many all‑in‑one plugins combine these steps.
- Insert an M/S encoder as the first effect on your stereo dialogue track. Ensure the plugin outputs separate Mid and Side signals (some call them “M”, “S” outputs).
- Route the outputs to two new aux tracks (e.g., MID and SIDE), or use the plugin’s internal routing if it has separate processing lanes.
- Solo the mid channel. Listen critically: you should hear only the centered dialogue (no stereo ambience). Apply EQ to remove low‑end rumble (below 80 Hz) and any muddy frequencies (200–400 Hz). Use a compressor with a moderate ratio (2:1 to 4:1) to even out level fluctuations. If needed, add de‑essing to tame sibilance.
- Now solo the side channel. You will hear a thin, phasey sound – that’s the directional information and room ambience. Apply a gentle high‑pass filter (around 200–300 Hz) to remove subsonic noise or rumbles that only exist in the difference channel. You can also add a little high‑frequency shelf boost to brighten the space without affecting the dialogue.
- Un‑solo both channels and adjust the relative level of the side fader. Start with the side level at the same fader position as the mid. To make the dialogue more intimate or focused, lower the side by 2–6 dB. To widen the scene (e.g., for an outdoor conversation with wind or city ambience), raise the side slightly – but be cautious: too much can make the dialogue sound hollow or phasey.
- Insert an M/S decoder after the two aux tracks (or use the encoder plugin’s decoder section) to sum the processed mid and side back into a normal stereo signal. Listen to the final result and compare to the original by bypassing the whole chain.
If your plugin (e.g., FabFilter Pro‑Q 3) processes Mid and Side internally, you can skip the separate routing and simply apply EQ in mid/side mode directly on the original track. The principle remains the same.
Using a Mid/Side Encoder vs. Separate Tracks
Both methods work, but separate tracks give you more flexibility to use any compressor, reverb, or limiter on each component. All‑in‑one plugins are faster and often have built‑in monitoring to solo mid or side. For dialogue purposes, the separate‑track approach is recommended until you are comfortable with the technique, because it lets you audition each component in isolation.
Advanced Techniques for Dialogue
Side‑Chain Compression from Mid to Side
This advanced trick reduces the stereo ambience whenever the dialogue becomes loud enough to trigger a compressor. Place a compressor on the side track and set its side‑chain input to listen to the mid track. The compressor will attenuate the side channel when the dialogue is present, leaving the background more spacious only during pauses. This helps the voice remain focused without losing all of the room tone. Use a medium attack (10–20 ms) and fast release (50–100 ms) to avoid pumping.
EQ Matching Between Mid and Side
If the ambient noise in the side channel has a resonant frequency that competes with the dialogue (e.g., a 250 Hz hum from an air conditioner), you can not notch it out of the side ONLY. Use a dynamic EQ set to mid/side mode, or a static EQ on the side track. This preserves the original character of the dialogue while cleaning the background.
Dynamic Mid/Side Processing
Some plugins like iZotope Neutron or FabFilter Pro‑Q 3 allow you to apply compression or EQ changes that react dynamically to the signal’s mid or side content. For example, you can automatically widen the image during quiet sections and narrow it during loud dialogue – a technique used in many broadcast productions to keep the voice consistent across varying stereo widths. Automating the side level with volume automation is the simplest, most transparent method.
Common Pitfalls and How to Avoid Them
Phase Cancellation Issues
Because the side channel is derived from a subtraction (L – R), if you apply heavy EQ boosts or all‑pass filters to it, the final stereo signal might develop phase cancellations that make the image sound unstable or comb‑filtered. Always check the correlation meter on your master bus. Values consistently below +0.3 indicate potential mono compatibility problems. If your correlation meter shows the signal dipping below 0 (out of phase), reduce side processing or adjust EQ curves until the meter returns to positive values.
Over‑widening Destroying Intelligibility
It’s tempting to boost the side for a dramatic, cinematic space. But for dialogue, excessive width can make the voice sound as if it’s coming from two speakers with a hole in the middle. The result: the brain has to work harder to locate the voice. Limit side boosts to 3‑6 dB relative to the original level. If you need more width, consider adding subtle stereo reverb to the side track instead.
Excessive Ambience Bleeding
If your original recording already has a lot of room sound (e.g., on‑set dialogue with reverberation), boosting the side will bring out that ambience even more. In such cases, reduce the side level or apply a gate/expander to the side channel to let only the directional information through. You can also compress the side heavily with a fast attack to reduce dynamic swings in the room sound.
Listening and Monitoring Checks
Mono Compatibility Check
After applying M/S processing, collapse the mix to mono with a plugin (like the “Mono” button on a utility gain plugin). The dialogue should remain at the same level and be clear. If the voice becomes quieter or hollow, the side processing is causing phase issues. Adjust the side EQ or reduce the level of the side channel until mono playback sounds natural.
Correlation Meter
Use a goniometer or correlation meter (available in many analyzers like iZotope Insight, Voxengo Span Plus, or basic DAW plugins). The correlation should stay positive, ideally between +0.5 and +1.0. If it dips below 0, you have a phase error. Read more about phase correlation on Sound On Sound.
Using Reference Tracks
Load a professionally mixed dialogue scene (from a movie or well‑produced podcast) into your DAW on a separate track. Use a mid/side analyzer (like the Flux Stereo Tool) to observe the mid/side balance of the reference. Typically, dialogue‑heavy scenes have a mid level 6‑12 dB higher than the side at most frequencies. Try to match that ratio in your mix for a natural, industry‑standard sound.
Conclusion
Mid/side processing gives you surgical control over the spatial placement of dialogue. By working on the mid and side components independently, you can center the voice, reduce problematic ambience, and widen or narrow the soundstage without compromising mono compatibility. The technique is used daily by broadcast engineers, film re‑recording mixers, and podcast producers. Start with subtle adjustments — around 2‑3 dB of side reduction or a simple low‑cut on the mid — and gradually explore more advanced concepts like mid/side compression and dynamic widening. With practice, you will develop an intuitive sense for how much width a dialogue track needs. For further reading, check out Pro Tools Expert’s guide to M/S processing and Mixing Lessons’ beginner tutorial.