# What Is Audio Mastering? A Guide for Voice Producers

URL: https://anyvoice.app/journal/what-is-audio-mastering-a-guide-for-voice-producers
Type: blog
Locale: en
Published: 2026-09-21
Updated: 2026-09-21

---

> What audio mastering is, how the signal chain works, platform LUFS targets, and what changes when your source audio is AI-generated.

What is audio mastering? It is the final processing stage between your mix and distribution: the step where tonal balance, dynamics control, and platform-compliant loudness come together before a file reaches a listener's ears.

## The one-sentence definition: and why it matters more than the phrase suggests

Audio mastering is the process of taking a finished stereo mix and preparing it for distribution. That sentence is technically accurate and almost useless on its own.

What mastering actually means in practice is applying a sequence of calibrated processing decisions: EQ, compression, stereo enhancement, limiting. Applied to a final mix file, these decisions meet three objectives simultaneously: tonal consistency, controlled dynamics, and platform-compliant loudness. A mastered track sounds coherent on a $15 pair of earbuds and a $4,000 pair of Genelec studio monitors. It meets the integrated LUFS target that Spotify or ACX requires without clipping or perceptible distortion. And it stays sonically aligned with the other titles on the same release or in the same game.

Mastering is not a fix for a bad mix. If the low end is muddy, the sibilance is aggressive, or the stereo image is narrow, mastering can address those problems at the margins, but it will not fix them at the source. The mastering engineer's headroom is limited by what the mix provides. For voice producers working with AI-generated audio, this distinction is especially practical: if the clone has synthesis artifacts at 8-10 kHz or uneven formant stability across takes, address those before the mastering stage.

## The mastering chain: what each stage does to the signal

A standard mastering chain runs in roughly this order, though the exact sequence depends on the material and the engineer's judgment.

**High-pass filter.** Removes sub-sonic rumble below 20-30 Hz. On voice material, energy below 80 Hz is rarely useful and often introduces low-end buildup that muddies the mix.

**Equalisation.** Broad, surgical adjustments to the full-frequency response. Typical moves are 0.5 to 2 dB, not the 6-10 dB corrections you might apply to an individual track during mixing. The goal is tonal balance: the low-mids don't feel congested, the high end has presence without harshness, the voice cuts through without being sharp.

**Stereo imaging.** Adjustments to width and mono compatibility. Most voice content: audiobooks, podcast narration, game dialogue, should stay largely mono-compatible. Check that critical dialogue collapses cleanly to mono before finalising.

**Compression.** Optical, VCA, and multiband compressors each affect transients and sustain differently. At the mastering stage, compression is typically gentle (2-4 dB gain reduction maximum) and aimed at gluing the mix together rather than shaping individual elements. On voice material, use a slow attack (30-80 ms) to preserve transients and a medium release (100-300 ms) to maintain naturalness.

**Limiting.** The final stage. A true peak limiter sets the ceiling, prevents inter-sample peaks that cause distortion during lossy encoding, and sets the integrated loudness target. Most limiting at mastering is 3-6 dB of gain reduction. More than 6 dB of limiting on voice typically introduces audible pumping and reduces intelligibility on narrow-bandwidth playback devices.

![Digital loudness meter showing LUFS readings in a professional mastering session](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/anyvoice/2026-09/25be84-inline1.webp)

## LUFS, true peak, and the platform targets that actually matter

LUFS stands for Loudness Units relative to Full Scale: a perceptual loudness measurement that accounts for how human hearing weighs different frequencies. Integrated LUFS measures average loudness over the full duration of a file. Momentary and short-term LUFS measure instantaneous loudness over 400 ms and 3-second windows respectively.

Each distribution platform specifies a target integrated LUFS level and normalises uploads to that value. Delivering louder than the target does not give you a loudness advantage: the platform turns it down. Delivering quieter wastes dynamic range.

Current platform targets:

- 
**Spotify:** -14 LUFS integrated, -1 dBTP true peak

- 
**Apple Music:** -16 LUFS integrated, -1 dBTP true peak

- 
**YouTube:** -14 LUFS integrated, -1 dBTP true peak

- 
**Tidal:** -14 LUFS integrated, -1 dBTP true peak

- 
**ACX (audiobooks):** -23 to -18 LUFS integrated, -3 dBTP true peak, -60 dBRMS noise floor

- 
**EBU R128 (broadcast):** -23 LUFS integrated, -1 dBTP true peak

*Session note: ACX standards are stricter than streaming music targets. The integrated loudness window is narrow (-23 to -18 LUFS, with most QC reviewers preferring -23 to -20) and noise floor requirements are tight (-60 dBRMS ambient). Check these specs before you master, not after. A file that passes Spotify normalisation can fail ACX QC and require a full re-master.*

For game dialogue, LUFS targets vary by engine and implementation. Unreal Engine 5's default audio architecture expects -23 LUFS for dialogue buses when using Dialogue Voice Attenuation. Unity targets depend on the game's dynamic range profile. If no project spec is provided, target -20 to -18 LUFS integrated and let the game engine's gain staging handle the final level.

## Mastering AI-generated voice: where synthetic signals behave differently

Most mastering guides assume the source is a live recorded take. AI-generated voice has different spectral and dynamic characteristics that affect how each stage of the mastering chain behaves.

**High-frequency synthesis artifacts.** AI voice synthesis often introduces subtle brightness inconsistencies or faint metallic artifacts in the 5-12 kHz range. These typically appear as a slight harshness absent from clean recorded takes. A gentle high-shelf cut (-1 to -1.5 dB above 8 kHz) often reduces this without softening presence. Test before and after on earbuds: synthesis artifacts are more audible on consumer headphones than on flat studio monitors.

![Frequency spectrum analysis and EQ curve visualization in a digital audio workstation](https://fdzlnqpwsaniezitwiuw.supabase.co/storage/v1/object/public/cms-media/anyvoice/2026-09/d686cb-inline2.webp)

**Formant consistency across takes.** When assembling multi-take sessions from an AI voice, different generation parameters or slight model drift can create formant shifts between consecutive takes. These show up as tonal inconsistency in the assembled file: one paragraph sounds slightly brighter or darker than the next. Catch these before mastering with a spectral comparison tool. Mastering EQ can smooth broad frequency imbalances but cannot correct take-to-take formant drift.

**Dynamic range.** AI-generated voice at default synthesis settings tends to have less micro-dynamic variation than a recorded performance, so the signal is perceptually flatter. This means gentle multiband compression at mastering may have less to do, but it also means the output can fatigue faster on extended listening. Consider preserving rather than compressing the small dynamic variations the model does generate.

**Noise floor.** AI voice renders typically have a near-silent noise floor (-90 dBFS or lower), unlike recorded takes which carry room tone at -55 to -65 dBRMS. For ACX delivery, this works in your favour: you meet the noise floor spec by default. For narrative content where the absence of room tone creates an unnatural feel, some producers add low-level room tone (-65 to -60 dBRMS) before mastering to increase perceived naturalness at normal listening volumes.

## The tools in a typical mastering chain

The mastering chain can run as a hardware outboard setup, a software-only DAW chain, or through an AI-assisted mastering service.

**Hardware outboard.** Tube and transformer-based analogue processors add harmonic character that some engineers prefer for music mastering. For voice and dialogue, the colour from analogue hardware is rarely the priority. Clarity, intelligibility, and accurate metering matter more. Analogue mastering is less common for voice work than for music.

**DAW plugin chains.** iZotope Ozone (full suite), FabFilter Pro-L2 (limiting), Nugen Audio MasterCheck (loudness metering), and Brainworx bx_masterdesk are common choices for professional software mastering chains. iZotope Ozone's AI-assisted modules: Master Assistant and EQ Match give a solid starting point that experienced engineers then refine manually.

**AI mastering services.** LANDR, CloudBounce, and eMastered offer automated mastering based on ML models. These services are fast and cost-effective for high-volume voice production: large NPC dialogue batches, multilingual audiobook editions. Quality varies. For music, AI mastering services are often criticised for generic EQ decisions. For voice content with a consistent spectral profile, results are more predictable. Run a batch test before committing a full project.

*Session note: AI mastering services use their own LUFS targets and true peak limits. Verify the output against your delivery spec rather than trusting the platform preset. A LANDR "streaming" preset does not guarantee ACX compliance.*

## Common mastering mistakes that cost voice clarity

**Over-limiting.** The most common error. A true peak limiter squeezing 8-10 dB of gain reduction to hit a loudness target does not increase perceived loudness on streaming platforms: normalisation brings it back down. What it does do is remove all dynamic variation from the signal and introduce distortion that shows up as harshness on consumer speakers.

**Mastering the mix bus, not a printed file.** Master from a printed stereo file: bounced offline with the mix bus processing bypassed, not from the live mix bus with plugins running. Mix bus processing applied during mastering and during the bounce creates accumulated decisions that are difficult to untangle.

**Using the mastering room as a substitute for accurate monitoring.** Mastering decisions made on a single playback system often fail to translate. Check the output at minimum on studio monitors (flat response), earbuds or consumer headphones, and a Bluetooth speaker. For voice content, intelligibility on a phone speaker at 70% volume is a reliable proxy test.

**Skipping the codec check.** Lossy encoding: MP3 at 128 kbps, OGG Vorbis for game audio, introduces inter-sample peaks and can amplify high-frequency artifacts. Run your mastered file through the delivery codec before final QC. What sounds clean at 24-bit WAV may have audible distortion at 128 kbps MP3 if the true peak limit was set at 0 dBTP instead of -1 dBTP.

## Before your next session

The practical checklist before mastering any voice session:

- 
Print the mix to a clean stereo file; bypass the mix bus

- 
Verify the delivery spec before touching the chain

- 
Check for synthesis artifacts in the 5-12 kHz range if the source is AI-generated

- 
Check mono compatibility, especially for content delivered on mobile devices

- 
Set the limiter ceiling to -1 dBTP minimum; use -3 dBTP for ACX delivery

- 
Encode to delivery format and run a final codec-check playback

Mastering is a short step in the production chain with a disproportionate impact on how the final audio holds up across listening contexts. For voice work: recorded, synthesised, or blended, getting the LUFS target, the EQ balance, and the limiter settings right is the difference between audio that sounds professional at delivery and audio that sounds like a draft.

## FAQ

### What is the difference between mixing and mastering?

Mixing balances individual tracks: adjusting levels, panning, EQ, and effects on each element until the session sounds coherent as a stereo file. Mastering takes that finished stereo file and prepares it for distribution: tonal balance, dynamics control, loudness calibration to platform specs, and codec-check before delivery. You mix multi-track sessions; you master the finished stereo print.

### What LUFS target should I use for audiobook mastering?

ACX (the Audible/Amazon distribution standard) requires -23 to -18 LUFS integrated loudness, with most QC reviewers preferring the tighter -23 to -20 LUFS window. True peak must be at or below -3 dBTP and the noise floor must be at -60 dBRMS or lower. These specs are stricter than streaming music targets: verify before you master, not after.

### Do AI mastering services work for voice content?

AI mastering services like LANDR and eMastered perform reliably on voice with a consistent spectral profile: the ML models make predictable EQ and limiting decisions on narrow-bandwidth material. They are less suited to complex multi-character dialogue or content with significant take-to-take variation. Always verify the output against your delivery spec rather than trusting the platform preset.

### How do you master AI-generated voice differently from recorded voice?

AI voice synthesis often introduces subtle high-frequency artifacts in the 5-12 kHz range that may need a gentle high-shelf cut at mastering. AI voice also has a near-silent noise floor versus room tone on recorded takes, which some producers compensate for by adding low-level room tone (-65 to -60 dBRMS) before mastering. Dynamic range is typically narrower on AI voice, so heavy multiband compression is rarely needed.

### What is true peak and why does it matter for audio mastering?

True peak measures the actual peak level of a waveform after digital-to-analogue conversion, including inter-sample peaks that standard peak metering doesn't detect. Lossy encoding formats (MP3, AAC, OGG) can amplify these inter-sample peaks above 0 dBFS, causing distortion in the decoded file. Setting a true peak ceiling of -1 dBTP at mastering prevents audible clipping on delivery.

### What loudness target should I use for game audio dialogue?

Unreal Engine 5's default dialogue audio architecture expects -23 LUFS for voice buses when using Dialogue Voice Attenuation. Unity targets depend on the game's dynamic range profile. In the absence of a project spec, target -20 to -18 LUFS integrated and let the game engine's gain staging handle the final level. Confirm the spec with the audio lead before mastering a full dialogue batch.

### Can I master inside the same DAW session I mixed in?

Technically yes, but it introduces risk. If the mix bus has active processing during the bounce, you are baking those decisions into the print and then applying mastering on top: accumulated processing that is difficult to separate later. Standard practice is to print a clean stereo file with the mix bus bypassed, open a fresh mastering session, and work from that print.