Captions, Audio & Rhythm5 min readLast updated September 2026

How to Clean Up Spoken Video

Handling pauses, filler utterances, repeated takes, false starts, and background noise floors.

Quick Takeaway

Clean up spoken video audio by: 1) Trimming dead silences, 2) Deleting false start takes, 3) Filtering verbal fillers, and 4) Applying voice enhancement EQ.

1. The order to do it in

Spoken audio cleanup has a sensible sequence, and following it saves repeating work. Structural cuts first, because there is no point polishing audio you are about to delete. Then the pacing pass that removes pauses, fillers and stutters. Then levels — the voice, then anything sitting underneath it. Captions last, because they are derived from the finished timeline.

  • Cut the sections and takes you are not using.
  • Remove pauses, filler words and stutters.
  • Set the voice level, then place music and effects underneath it.
  • Generate captions once the timeline is settled.

2. Removing bad takes

When a sentence was attempted three times, only one of them belongs in the video — usually the last, because by then you knew what you were trying to say. These are easy to find because they are language-shaped: the same phrase appearing twice in a row, an abandoned clause, a restart after a pause. Removing them is part of the standard cleanup pass rather than a separate job.

Take selection by description
"I restart that sentence a few times — keep the last attempt and cut the others."

3. Levels and ducking

Volume in an edit is relative, not absolute: an adjustment is a change to a clip's own level rather than a fixed loudness. The two things worth getting right are a consistent voice level across the whole video, and anything underneath the voice staying underneath it. A music bed that occasionally rises over a sentence is more distracting than one that is simply too loud throughout, because the viewer notices the movement.

  • One consistent level for the speech, so the viewer never adjusts their volume.
  • Music and effects clearly below it, not merely quieter.
  • A dip under a specific line is a curve — it falls, holds, and returns rather than staying down.
  • Fades of around a fifth of a second sound natural; much faster reads as abrupt.
A dip, not a permanent drop
"Duck the music under my voice between 0:12 and 0:19 and bring it back after."

4. Music under speech

Background music on spoken content is a support, not a feature. It should be quiet enough that you stop noticing it and would only notice its absence, and it should cover the whole video rather than stopping partway through. If a bed is shorter than the edit, it needs to loop; music that ends halfway through a video is a defect, not a stylistic choice.

  • Instrumental only. Lyrics compete directly with speech.
  • Set it to run the whole length of the edit, looping if the track is shorter.
  • Fade it out over the last few seconds rather than cutting it dead.
  • If in doubt, quieter. Nobody has ever complained that background music was too subtle.

5. What cannot be fixed in post

It is worth knowing the limits so you spend your effort in the right place. Editing removes things and changes levels; it does not un-record a room. Heavy echo, a microphone that was too far away, wind noise, or a recording that clipped into distortion are all recording problems, and the honest fix is to record that part again rather than to spend an hour trying to rescue it.

  • Room echo: a property of the space, not the file.
  • Clipping and distortion: the information is gone.
  • Distant, thin microphone sound: no amount of level adjustment creates presence.
  • Do fix in post: pacing, filler, take selection, levels, balance between sources.
Tip

The cheapest audio upgrade is not a plugin, it is getting the microphone closer and recording somewhere with soft surfaces.

Common Mistakes to Avoid
✖ Setting music level once and never checking it

Why it fails: A bed that works under a loud section will swamp a quiet one, and the viewer hears the imbalance.

✔ Better approach: Duck the music under the sections where the voice drops, rather than picking one compromise level.

✖ Using a music track shorter than the video

Why it fails: The bed stops partway through and the sudden silence is more noticeable than the music ever was.

✔ Better approach: Set the music to cover the full length of the edit and loop if needed.

✖ Polishing audio before locking the cut

Why it fails: Every structural change moves the material you just balanced.

✔ Better approach: Cut first, then set levels, then caption.

Frequently Asked Questions

Common questions around this editing workflow.

Quiet enough that you have to listen for it. If you can follow the speech comfortably without concentrating, the level is right.

Want to edit this way without doing everything manually?

Try Cubix conversational AI video editor free for 7 days.