sound-design-and-mixing
Best Practices for Exporting Dialogue Mixes for Different Platforms
Table of Contents
Why Platform-Specific Dialogue Mixing Matters
Exporting dialogue mixes is one of the most critical steps in audio post-production, yet it is often treated as an afterthought. Dialogue is the primary carrier of narrative and information in podcasts, films, broadcasts, and video content. If your mix sounds clear and natural on your studio monitors but falls apart on a listener’s phone speaker, car stereo, or streaming service, the entire production loses value. Each distribution platform imposes unique technical constraints — loudness targets, codec behavior, playback environments, and listener expectations — that demand a tailored exporting approach. Understanding and applying best practices for platform-specific exporting ensures your dialogue maintains its intended clarity, presence, and emotional impact across every playback scenario.
This article provides a production-ready framework for exporting dialogue mixes for different platforms. We will cover core technical specifications, advanced mixing techniques that directly improve export quality, platform-specific strategies for streaming, broadcast, cinema, gaming, and mobile, and common pitfalls to avoid. By the end, you will have a repeatable workflow that saves time, reduces QC failures, and delivers professional results every time.
Understanding Platform Audio Specifications
Before you export a single file, you must know the destination platform’s requirements. These specifications are not arbitrary; they exist to ensure consistent playback across millions of devices and listening environments. Ignoring them leads to rejected deliveries, distorted audio, or a poor listener experience.
Loudness Standards
Loudness is measured in LUFS (Loudness Units relative to Full Scale) and is the single most important metric for dialogue exports. Different platforms enforce different targets. The ITU-R BS.1770 standard forms the basis for most modern loudness measurement, and you should calibrate your meters to this specification.
- Podcasts and streaming services (Spotify, Apple Podcasts, YouTube, Amazon Music): Target an integrated loudness of -16 LUFS ±1 LU, with a true peak maximum of -1 dBTP. This standard provides consistent loudness across episodes and aligns with the loudness normalization applied by most streaming players. Some platforms like YouTube may also normalize content to -14 LUFS, but -16 LUFS remains the safest target for podcast hosts.
- Broadcast television (ATSC A/85 in North America, EBU R128 in Europe, ITU-R BS.1770): Target -23 LUFS ±0.5 LU for dialogue. Broadcast specifications also include loudness range (LRA) limits and short-term loudness constraints to prevent excessive dynamic swings during commercials or dramatic scenes. For EBU, the maximum short-term loudness is -23 LUFS, and true peak must be ≤ -2 dBTP; for ATSC, true peak is often ≤ -6 dBTP.
- Cinema (Dolby Atmos, DCP): Cinema mixes are typically calibrated to -20 dBFS pink noise at the mixing stage, with dialogue averaging around -27 dBFS to -31 dBFS on a standard SPL meter. For digital cinema packages, follow SMPTE ST 2094 and Dolby’s metadata guidelines. The dynamic range is much wider; dialogue must remain intelligible against loud effects while preserving emotional nuance.
- Gaming and interactive media: Game audio engines (Wwise, FMOD) handle real-time mixing, but exported dialogue assets should target -18 dBFS to -24 dBFS average with ample headroom for in-game dynamic mixing. Peak levels should not exceed -3 dBFS to allow engine processing without clipping.
Always verify the current specifications on the platform’s official documentation. Standards evolve, and what was correct last year may have changed. For a comprehensive reference, consult the ITU-R BS.1770 loudness standard and the EBU R128 specification.
Sample Rate and Bit Depth
Choosing the correct sample rate and bit depth preserves audio fidelity and prevents resampling artifacts during encoding. The Nyquist theorem dictates that the sample rate must be at least twice the highest frequency you wish to capture. For dialogue, the audible range up to 20 kHz is well covered by 44.1 kHz and 48 kHz.
- Sample rate: 44.1 kHz is standard for music and podcasts (CD quality). 48 kHz is standard for film, television, and video, as it aligns with video frame rates. 96 kHz is occasionally used for high-resolution archival or Dolby Atmos renders but is unnecessary for final delivery in most cases. Export at the rate native to your project to avoid sample rate conversion that can introduce aliasing or timing errors.
- Bit depth: 24-bit is the industry standard for dialogue mixing and mastering because it offers 144 dB of dynamic range, providing ample headroom for processing and fader moves. Export masters at 24-bit, then dither to 16-bit only if the destination explicitly requires it (e.g., CD or some podcast hosts that expect 16-bit WAV). Avoid exporting at 16-bit unless forced, as the reduced dynamic range can increase quantization noise and degrade quiet dialogue passages. When dithering, use noise-shaped dither for 16-bit exports to push noise into less audible high frequencies.
File Format and Codec Strategy
The file format you choose affects file size, metadata support, and compatibility with the distribution platform’s encoding pipeline. Always start with a high-quality master and encode to lossy formats only as needed.
- WAV (uncompressed): Best for master archives, broadcast delivery, and any workflow where quality must be pristine. Use WAV for final masters before encoding into lossy formats. Broadcast WAV (BWF) includes metadata like timecode, which is essential for television and cinema.
- FLAC (lossless compressed): Suitable for archiving with reduced file size, but not universally supported by all delivery platforms. FLAC is excellent for personal archives but avoid it for direct delivery to most streaming services.
- MP3 (lossy): Still required by some older podcast hosts and low-bandwidth platforms. Export at 320 kbps (CBR) or at least 256 kbps for dialogue to minimize audible artifacts. For speech-only content, a lower bitrate (128 kbps) may be acceptable, but always test. Be aware that MP3 introduces artifacts like pre-echo and spectral holes that can degrade sibilance.
- AAC (lossy, .m4a): Preferred by Apple Podcasts, YouTube, and many streaming services. Export at 256 kbps or higher for high-quality dialogue. AAC generally outperforms MP3 at equivalent bitrates, especially for speech. Opus (used by YouTube for some streams) offers even better efficiency, but delivery formats rarely include Opus directly.
When preparing dialogue mixes for platforms that re-encode your file (e.g., Spotify, YouTube, most podcast hosts), start with a 24-bit WAV master at the correct sample rate. The platform’s encoder will do a better job if you give it a clean, uncompressed source than if you feed it a lossy file already degraded by prior encoding. Some platforms also accept FLAC for lossless delivery; check their guidelines.
Dialogue Mixing Techniques That Improve Export Quality
Export settings are only half the equation. The quality of your dialogue mix before export determines how well it survives codec compression and acoustic playback variations. The following techniques are essential for creating a mix that translates reliably across platforms.
Gain Staging and Level Consistency
Dialogue levels should be consistent within a scene, across scenes, and throughout the entire program. Use clip gain or volume automation to even out performance variations before applying any dynamics processing. Clip gain is preferable for initial leveling because it happens before any inserts, allowing compressors to see a consistent input level. Aim for dialogue to sit between -12 dBFS and -6 dBFS on your meters during normal speech, with peaks reaching no higher than -3 dBFS. This headroom gives your compressor, limiter, and final output stage room to operate cleanly without hitting the ceiling. If you are mixing for broadcast, leave even more headroom (peaks around -6 dBFS) to accommodate later loudness adjustments.
Dynamic Range Compression
Compression is the dialogue mixer’s primary tool for controlling level variation and improving intelligibility. For platform-specific exports, you may need to adjust compression depth and style. Consider using multiband compression for dialogue if you need to control specific frequency ranges independently — for example, taming low-mid rumble without squashing highs.
- For podcast and streaming: Use moderate-to-heavy compression (ratio 2:1 to 4:1) with a fast attack (10–30 ms) and medium release (40–80 ms). Aim for 3–6 dB of gain reduction on average. This creates a dense, upfront sound that works well on earbuds and car speakers. A slow attack can preserve consonants like plosives, while fast attack helps control sibilance.
- For broadcast and cinema: Use lighter compression (ratio 1.5:1 to 2.5:1) with slightly slower attack (20–50 ms) to preserve dynamic nuance. Heavy compression can make dialogue sound fatiguing in a theater or over a long broadcast run. For cinema, consider using parallel compression to retain transients while raising average level.
- For gaming and mobile: Use aggressive compression (ratio 4:1 or higher) with a fast attack (5–15 ms) and short release (20–40 ms). Game dialogue must remain intelligible in noisy, unpredictable playback environments. However, avoid pumping artifacts that might be exposed when the game audio changes rapidly.
Always use a high-quality compressor with low distortion. Consider using a dedicated dialogue processor like the Waves CLA-2A, FabFilter Pro-C 2, or iZotope RX Dialogue Leveler for more transparent results. For speech, look for plug-ins that offer a “mastering” mode with soft-knee characteristics.
Equalization for Clarity and Compatibility
EQ shapes the tonal balance of dialogue and directly impacts how it reproduces on different playback systems. Use surgical EQ cuts before broad boosts to maintain phase coherence.
- High-pass filter: Roll off frequencies below 80 Hz to remove rumble, HVAC noise, and proximity effect. Use a steeper slope (24 dB/octave) for aggressive cleaning, but be careful not to remove natural chest resonance. For deep male voices, a 60 Hz cutoff may preserve warmth; for female voices, 80–100 Hz is often safe. Use a linear-phase EQ for high-pass filters if you are concerned about phase shift across the frequency range.
- Presence and articulation boost: A gentle shelf boost (2–4 dB) around 2–5 kHz increases sibilance, fricatives, and consonant articulation — all critical for intelligibility on small speakers and in noisy environments. Avoid over-boosting, which can cause sibilance to become harsh, especially after lossy encoding. Consider a gentle bell boost at 3 kHz for airy clarity.
- Low-mid control: The 200–500 Hz range often causes boxiness or muddiness. A gentle cut (2–3 dB) in this region can clean up the dialogue and improve clarity. Listen for nasality around 1 kHz and apply a narrow cut if needed.
- High-frequency roll-off: A gentle low-pass filter above 15 kHz reduces digital artifacts and hiss without affecting speech bandwidth. This is especially beneficial before lossy encoding, which can add high-frequency noise. For older voice actors with a lot of sibilance, a subtle de-esser before EQ is often more effective than EQ alone.
True Peak Limiting and Headroom
True peak limiters prevent inter-sample peaks from causing distortion when the audio is decoded by consumer DACs. True peak is higher than sample peak and can be up to 3 dB above the sample peak for some signals. For all dialogue exports, set your limiter to a ceiling of -1 dBTP (or -2 dBTP if the platform specifies). Use a transparent limiter with a fast release (1–5 ms) and moderate gain reduction (2–4 dB) to catch stray peaks without pumping. Avoid pushing dialogue too close to the ceiling; leaving 1–2 dB of headroom will give downstream encoders room to work without forcing them to clip. Some limiters offer a look-ahead function that prevents overshoots; enable it for safer exports.
Loudness Metering and Alignment
You cannot export dialogue to platform specifications without proper metering. Use a loudness meter that supports ITU-R BS.1770-4 and can display integrated, short-term, and momentary loudness, as well as true peak. Many meters also show loudness range (LRA) and can be gated to exclude silence. Gating is important for accurate integrated loudness measurement because long pauses can lower the average artificially.
- Measure the integrated loudness of your entire program (or at least a representative section). Use gating (usually -70 dBFS threshold) to exclude silences longer than 1 second.
- If the loudness is above the target (e.g., -12 LUFS when you need -16 LUFS), reduce gain on the entire mix by the difference (4 dB) and re-measure. Do not rely on the limiter to achieve loudness reduction; gain staging is cleaner.
- If the loudness is below the target, apply gain staging or increase compression to bring it up, then re-measure. Alternatively, use a “loudness maximizer” plug-in designed to raise integrated loudness without distortion.
- Check short-term and momentary loudness to ensure no section exceeds platform limits (e.g., broadcast often requires that no 3-second segment exceed -23 LUFS). Use the meter’s history graph to spot problematic sections.
- Verify true peak does not exceed the limit. If it does, reduce the mix level or adjust the limiter ceiling.
Repeat this process for each platform target. You can create different export presets in your DAW to streamline the workflow. Some audio editors like iZotope RX or Adobe Audition have batch loudness processing tools that can automatically adjust multiple files to a target.
Platform-Specific Export Strategies
While the principles above apply broadly, each platform family has unique considerations that demand targeted adjustments.
Podcast and On-Demand Streaming
Podcasts and streaming services use loudness normalization, meaning the platform adjusts playback level to a consistent loudness (usually -16 LUFS). If your mix is already at -16 LUFS, it will play at unity gain. If it is louder, the platform will attenuate it, potentially causing your dialogue to sound quieter than the normalized ads or music. If it is quieter, the platform will boost it, which can raise noise floors and background artifacts. Some platforms like YouTube also apply a gain boost to content below -14 LUFS, so you may want to target exactly -14 LUFS for video content.
Export your podcast master as a 24-bit 44.1 kHz WAV file at -16 LUFS integrated, with true peaks no higher than -1 dBTP. Use a moderate compression ratio (3:1 to 4:1) to maintain consistent level and a gentle limiter to catch peaks. Avoid excessive noise reduction that leaves the dialogue sounding sterile or gated. Listeners value naturalness even more than absolute silence. For stereo podcasts, keep the dialogue centered and use wide-panned music beds carefully to avoid masking.
For platforms like Spotify and Apple Podcasts, refer to their official guidelines: Spotify Podcast Technical Specifications and Apple Podcasts Mastering Guidelines for detailed requirements. Spotify also recommends including a few seconds of silence at the end to avoid abrupt cuts.
Broadcast Television
Broadcast specifications are stricter than streaming. The EBU R128 and ATSC A/85 standards require that dialogue loudness average -23 LUFS ±0.5 LU, with a maximum short-term loudness of -23 LUFS and a true peak limit of -2 dBTP (EBU) or -6 dBTP (ATSC). Broadcast is unforgiving — a non-compliant mix can be rejected or cause automatic level shifting during transmission. Additionally, many broadcasters require metadata (dialnorm parameter) embedded in the audio stream to inform the receiver’s normalization.
When exporting for broadcast, use a lighter compression ratio (2:1) and maintain more dynamic headroom in the raw mix. Apply a true peak limiter with a ceiling of -2 dBTP (EBU) or -6 dBTP (ATSC) — set the limiter accordingly. Always export as a 24-bit 48 kHz WAV file. Use metadata to label the audio type (dialogue, music, effects) if the platform supports it. Test your mix on a loudness meter that supports the relevant broadcast standard before delivery. For ATSC, you may need to use a “dash” profile that measures loudness over the entire program, not just a segment.
Cinema and Theatrical
Cinema mixes require a different workflow. Dialogue is mixed to a reference level of -20 dBFS pink noise, with dialogue typically peaking around -10 dBFS and averaging -27 dBFS to -31 dBFS on the stage. The mix must handle the wide dynamic range of a theater environment: loud explosions in an action film can reach -5 dBFS or higher, while quiet dialogue may be -40 dBFS. The export format is often a DCP (Digital Cinema Package) using uncompressed PCM audio in BWF format, with sample rates of 48 kHz or 96 kHz, 24-bit. For Dolby Atmos, the delivery is an ADM BWF file containing object audio metadata.
For cinema dialogue, minimize compression to preserve dynamic impact. Use a true peak limiter only to prevent clipping during the loudest moments, with a ceiling of -1 dBTP. If you are mixing for Dolby Atmos, work within the Atmos renderer and follow Dolby’s delivery specifications. Dialogue is often placed in the center speaker with no or little spatial processing. Use a high-pass filter set to 60 Hz for dialog to avoid interfering with the LFE channel (which handles frequencies below 80 Hz).
Gaming and Interactive Media
Game dialogue assets must be exported as individual files (WAV, OGG, or ADPCM) with consistent gain staging so the game engine can mix them dynamically. Target an average level of -18 dBFS to -24 dBFS for dialogue, leaving plenty of headroom for the engine’s volume automation, environmental effects, and adaptive mixing. Use moderate compression (3:1 to 4:1) to keep the dialogue consistent but avoid aggressive limiting that might sound unnatural when repeated or variably pitched. Export at 48 kHz, 16-bit or 24-bit, depending on the engine’s capacity and platform requirements. Many game engines also require mono files for dialogue to allow spatial audio processing.
Game dialogue must also handle multiple playback environments: headphones, TV speakers, soundbars, and portable speakers. Test your dialogue at low playback volume and in simulated noisy conditions. Consider including a “mix minus” version (dialogue without stereo effects) for spatial audio engines that need a mono source. For virtual reality, ensure dialogue remains phase coherent for headphone playback with HRTF processing.
Mobile and Social Media Platforms
Instagram Reels, TikTok, YouTube Shorts, and other vertical video platforms compress audio aggressively. Dialogue exported for these platforms should be more heavily compressed (4:1 to 6:1) with a ceiling of -2 dBTP to survive the encoding. Use a presence boost around 3–5 kHz to cut through the limited bandwidth of mobile speakers. Export as AAC (256 kbps) or MP4 with a video container, at 44.1 kHz sample rate and 16-bit bit depth, unless the platform requires higher quality. For TikTok, note that audio is often processed with additional loudness normalization to -14 LUFS, so ensure your mix is consistent with that target.
Common Pitfalls and How to Avoid Them
Even experienced mixers make mistakes that degrade dialogue quality when exporting. Here are the most frequent issues and their solutions.
Inconsistent Loudness Across Episodes
If you produce a series, each episode must match the same loudness target. A listener who adjusts volume for one episode should not have to readjust for the next. Use the same export preset, the same limiter ceiling, and the same metering tool for every episode. Create a reference episode that you A/B against before finalizing any new mix. Inconsistent loudness is especially common when different mixers work on different episodes; establish a master template and enforce strict QC.
Over-Limiting for Loudness
Pushing your limiter to squeeze every last dB of loudness is a common mistake that destroys dialogue dynamics and introduces distortion. Dialogue does not need to be as loud as modern music. Respect the platform’s loudness target — if the target is -16 LUFS, do not try to hit -12 LUFS and hope the platform reduces it. You will lose nuance and increase listening fatigue. Over-limiting also raises the noise floor of background ambiences, making the dialogue sound artificial.
Sibilance De-essing and Lossy Encoding Issues
Sibilant consonants (s, t, sh, ch) can become harsh when compressed and then further encoded with lossy codecs. Apply a gentle de-esser (2–4 dB reduction in the 5–8 kHz range) before your compressor and limiter. The de-esser should reduce sibilance before the compressor has a chance to amplify it. After de-essing, check the export on consumer devices — if sibilance is still too prominent, apply a second de-esser stage after compression. Lossy codecs like MP3 and AAC introduce “swishing” artifacts on sibilance; test your mix by encoding to a low bitrate MP3 and compare to the original.
Ignoring Phase Correlation
Dialogue should always be mono-compatible (centered and in phase). If your dialogue mix has stereo effects, reverb returns, or spatial width that introduces phase issues, the dialogue may cancel out when summed to mono on some playback systems (e.g., many mobile phones and Bluetooth speakers). Use a phase correlation meter to ensure your mix is mono-compatible. If necessary, collapse the dialogue to a centered mono send in your export. For stereo dialogue (e.g., two characters on different channels), ensure they are in phase and summed correctly. Use a mid-side decoder to check that the side channel does not contain essential dialogue information.
Skipping the QC Playback Test
Do not rely solely on meters. Before final delivery, play your exported file on a minimum of three devices: studio headphones, a consumer Bluetooth speaker, and a laptop or phone speaker. Listen for distortion, pumping, sibilance, and intelligibility at low volume. If the dialogue sounds clear and natural on all three devices, your export is ready. Also check at -20 dB below normal listening level to simulate soft playback in a noisy environment.
Using the Wrong Dither Settings
When exporting to 16-bit, use appropriate dither (noise shaping recommended) and avoid applying dither multiple times. If you export a 24-bit file and then convert it again to 16-bit elsewhere, you may get double dither artifacts. Keep dithering to the final export stage.
Building an Efficient Export Workflow
To export dialogue mixes consistently and quickly, build a repeatable workflow in your DAW or audio mastering suite:
- Create platform presets: In your limiter and loudness meter, save presets for each platform target (e.g., “Podcast -16 LUFS,” “Broadcast -23 LUFS,” “Cinema -27 LUFS”). Include specific true peak limits and sample rate settings.
- Use a dedicated export template: Set up a master track with your limiter, loudness meter, and true peak meter. Route all dialogue tracks through this master. Save a separate template for each platform if needed.
- Normalize at the end: Adjust the master fader to achieve the target loudness rather than compressing or limiting more aggressively than needed. Use a gain stage before the limiter for coarse adjustment.
- Automate leveling: Use a dialogue leveler (like iZotope RX Dialogue Leveler or Waves Vocal Rider) to pre-level the dialogue before it hits the compressor, reducing the workload on the limiter. This also helps maintain consistent levels across various recording takes.
- Document your settings: Keep a log of platform specifications, target loudness, true peak limit, and any platform-specific EQ. When a guideline changes, update your preset immediately. Use version control for your presets to avoid confusion.
- Batch export when possible: For series, use batch processing tools to apply the same loudness chain to all episodes, then check each one individually for anomalies.
A well-organized workflow reduces the risk of human error and ensures every export meets the platform’s requirements without requiring last-minute fixes.
Final Quality Assurance Checklist
Before you deliver any dialogue export, run through this checklist:
- Loudness: Integrated loudness matches the target (±0.5 LU).
- True peak: No peak exceeds the platform’s limit (-1 dBTP for streaming, -2 dBTP for EBU broadcast, -6 dBTP for ATSC broadcast).
- Sample rate and bit depth: Match the project’s native settings and delivery requirements. Verify using metadata inspector.
- File format: WAV for masters, correct lossy format for delivery. Ensure extension matches platform expectations.
- Dialogue intelligibility: Listen on at least three devices, including a mobile speaker. Pay attention to sibilance and low-frequency clarity.
- No distortion or pumping: Check for artifacts from over-compression or limiting. Use a spectrogram to spot harmonic distortion.
- Mono compatibility: Verify dialogue remains centered and phase coherent. Sum to mono and listen for level loss or comb filtering.
- Version labeling: Clearly mark which platform the export is for in the filename (e.g., “Episode_10_Spotify_FINAL.wav”). Include version number and date.
- Silence handling: Ensure no unwanted DC offset or clicks at the beginning and end of the file. Trim silence to the nearest zero crossing.
- Metadata: If required, embed loudness metadata (dialnorm, program type) in the BWF file. Use tools like AudioMover or MetaData for broadcast.
Exporting dialogue mixes for different platforms is not a one-size-fits-all process. Each platform demands specific technical parameters, and the best mixers adapt their compression, EQ, loudness, and file formats to match. By understanding the platform requirements, applying targeted mixing techniques, and following a rigorous quality assurance workflow, you can ensure your dialogue sounds professional and intelligible wherever it is heard. Invest the time in setting up presets, testing your exports, and staying current with platform guidelines — your audience will notice the difference in clarity, consistency, and overall listening experience.
For further reading on loudness standards across platforms, refer to the Dolby Atmos delivery specifications and the AES technical documents on audio metering.