1. The order to do it in
Spoken audio cleanup has a sensible sequence, and following it saves repeating work. Structural cuts first, because there is no point polishing audio you are about to delete. Then the pacing pass that removes pauses, fillers and stutters. Then levels — the voice, then anything sitting underneath it. Captions last, because they are derived from the finished timeline.
- Cut the sections and takes you are not using.
- Remove pauses, filler words and stutters.
- Set the voice level, then place music and effects underneath it.
- Generate captions once the timeline is settled.
2. Removing bad takes
When a sentence was attempted three times, only one of them belongs in the video — usually the last, because by then you knew what you were trying to say. These are easy to find because they are language-shaped: the same phrase appearing twice in a row, an abandoned clause, a restart after a pause. Removing them is part of the standard cleanup pass rather than a separate job.
3. Levels and ducking
Volume in an edit is relative, not absolute: an adjustment is a change to a clip's own level rather than a fixed loudness. The two things worth getting right are a consistent voice level across the whole video, and anything underneath the voice staying underneath it. A music bed that occasionally rises over a sentence is more distracting than one that is simply too loud throughout, because the viewer notices the movement.
- One consistent level for the speech, so the viewer never adjusts their volume.
- Music and effects clearly below it, not merely quieter.
- A dip under a specific line is a curve — it falls, holds, and returns rather than staying down.
- Fades of around a fifth of a second sound natural; much faster reads as abrupt.
4. Music under speech
Background music on spoken content is a support, not a feature. It should be quiet enough that you stop noticing it and would only notice its absence, and it should cover the whole video rather than stopping partway through. If a bed is shorter than the edit, it needs to loop; music that ends halfway through a video is a defect, not a stylistic choice.
- Instrumental only. Lyrics compete directly with speech.
- Set it to run the whole length of the edit, looping if the track is shorter.
- Fade it out over the last few seconds rather than cutting it dead.
- If in doubt, quieter. Nobody has ever complained that background music was too subtle.
5. What cannot be fixed in post
It is worth knowing the limits so you spend your effort in the right place. Editing removes things and changes levels; it does not un-record a room. Heavy echo, a microphone that was too far away, wind noise, or a recording that clipped into distortion are all recording problems, and the honest fix is to record that part again rather than to spend an hour trying to rescue it.
- Room echo: a property of the space, not the file.
- Clipping and distortion: the information is gone.
- Distant, thin microphone sound: no amount of level adjustment creates presence.
- Do fix in post: pacing, filler, take selection, levels, balance between sources.
The cheapest audio upgrade is not a plugin, it is getting the microphone closer and recording somewhere with soft surfaces.
✖ Setting music level once and never checking it
Why it fails: A bed that works under a loud section will swamp a quiet one, and the viewer hears the imbalance.
✔ Better approach: Duck the music under the sections where the voice drops, rather than picking one compromise level.
✖ Using a music track shorter than the video
Why it fails: The bed stops partway through and the sudden silence is more noticeable than the music ever was.
✔ Better approach: Set the music to cover the full length of the edit and loop if needed.
✖ Polishing audio before locking the cut
Why it fails: Every structural change moves the material you just balanced.
✔ Better approach: Cut first, then set levels, then caption.
Frequently Asked Questions
Common questions around this editing workflow.