1. Outcomes vs mechanics
A traditional timeline asks you to perform an edit: pick the blade, split the clip, delete the gap, ripple the rest, add the transition. Natural language editing asks you to describe the finished state instead, and lets the editor work out which operations get there. The shift matters most on the repetitive work — pauses, captions, reframing, cutaways — where the mechanics are well understood and the only real decision is what you want.
- Mechanical: 'split at 00:42, delete to 00:47, ripple close the gap.'
- Outcome: 'cut the bit where I lose my train of thought.'
- Both are valid. The second one survives the footage changing underneath it.
2. How to point at a moment
You have three ways to identify a moment, and they are not interchangeable. Timestamps are exact and best when you have just watched the video. Quoted speech is the most reliable for spoken content, because it resolves against a word-level transcript. Visual descriptions work when the moment has no words attached, and they require the editor to look at the picture rather than the transcript.
- By time: 'trim from 01:12 to 01:45.'
- By what was said: 'cut everything before I say let me show you.'
- By topic: 'find the section about pricing and tighten it.'
- By what is on screen: 'speed up the part where the page is loading.'
Quoting a distinctive phrase is usually the most precise instruction you can give. 'The part where I say we grew 3x' beats 'around the middle' every time.
3. What to specify, and what to leave open
Specify anything that is a decision only you can make: the platform, the length, a moment that must stay in, a treatment you do not want. Leave open anything that is a craft judgement — where the cuts fall, which look suits the footage, whether a moment deserves a zoom. Over-specifying craft tends to make a video worse, because it forces a generic choice onto footage the editor can actually see.
- Worth saying: platform, duration, aspect ratio, a required moment, a banned treatment, brand colours.
- Usually better left open: cut rhythm, transition types, grade, where zooms land.
- Always worth saying: a number you actually need. 'Five clips' and 'under 30 seconds each' are binding when you say them.
4. Correcting without starting over
The conversation carries context, so a follow-up is understood against the edit that already exists. Adjusting a pass is different from running it again: re-running doubles the effect, while asking for less of it tunes what is already applied. If a whole reply went the wrong way, one undo takes back everything that reply did.
✖ Writing a prompt that describes the interface
Why it fails: Instructions like 'drag the clip to track 2' assume a manual workflow and say nothing about the intended result.
✔ Better approach: Describe the result: 'put the B-roll over me while I explain the pricing.'
✖ Bundling six unrelated requests into one sentence
Why it fails: If one of them is wrong you lose all six to a single undo, and it is harder to tell which instruction caused what.
✔ Better approach: Group related edits, review, then continue. Two or three outcomes per message is a good rhythm.
Frequently Asked Questions
Common questions around this editing workflow.