Voice ducking: music that knows when to get out of the way
7cubit Team
More in Podcasting

Background music works until it competes with the person speaking. A fixed music level can make a clear answer hard to follow, especially when the host gets quieter or a guest leans away from the microphone. Riding the music fader by hand solves the problem, but it also asks you to edit every entrance and exit.
Voice ducking handles that relationship for you: music drops while dialogue is present, then returns when the speech stops. The aim is not silence. It is a bed that supports the conversation rather than asking the listener to choose between two tracks.
How automatic ducking works in Podcast Studio
Add the music track to the Podcast Studio timeline and enable ducking. The system analyzes the host and guest dialogue tracks, lowers the music when a voice appears, and brings it back toward its original level when the speaking ends.
This is useful for an intro, a transition, or a low music bed under conversation. It also saves work when the spoken edit changes. If you cut filler words or rearrange a section, the music can follow the new timing instead of leaving old volume automation behind.
Set the balance for the show you are making
Podcast Studio provides a ducking threshold and a gain-reduction amount:
- Threshold: how loud speech must be before ducking begins. If background noise is triggering the response, adjust the threshold so ordinary mic noise does not make the music jump.
- Gain reduction: how far the music falls when someone speaks. More reduction prioritizes intelligibility; less reduction keeps the music more present for a higher-energy format.
These settings apply to the music track across the episode. Start with the default behavior, then listen to a quiet answer and a more animated exchange. A setting that works for an energetic intro may feel too aggressive under a careful interview.
Do not let the music trigger the mix
Voice ducking works from the dialogue tracks, so the source still matters. If the microphone carries a lot of room noise, breathing, or another speaker bleeding into the channel, listen for moments when the music drops for the wrong reason. Clean up the source or adjust the threshold before lowering the music more than the show needs.
Also check the return. Music that rises too quickly after every short phrase can sound nervous, while music that stays low for a long time can disappear from the arrangement. A bed should have a role between sentences, not only a volume change that proves the feature is active.
When ducking is a good fit
Use it when the music is meant to stay underneath speech: an opening theme that continues into the host’s welcome, a sponsor transition, or a light bed under commentary. It is especially useful when you are working quickly and do not need a hand-drawn volume envelope for every sentence.
Leave it off when the music is the content. A song that is meant to be heard in full should not be treated as a background bed. Also listen for sections where the host and music overlap by design; a sharp dip may fight the arrangement instead of helping it.
A simple review pass
- Listen to the first speech entry after the music starts.
- Check a quiet guest and a loud host to see whether the threshold behaves consistently.
- Listen to the return of the music after a sentence ends.
- Review any section you edited after enabling ducking.
Automatic does not mean invisible. The best setting is the one the listener stops noticing because every word is easy to follow.
The best music bed is felt before it is noticed. Let the voice own the sentence.
Voice ducking is a practical alternative to constant fader work, not a substitute for a good level or a thoughtful arrangement. Add the bed, set the response, and judge the result by the conversation.