1. What the Editor Reads First
An AI video editor does not treat your upload as a single opaque file. Before it can act on an instruction, it builds a few parallel readings of the same footage, and each one answers a different kind of question. A request about what was said is answered from the speech layer; a request about what was on screen has to be answered by actually looking at the picture.
- A word-level transcript, so a phrase like 'the part where I talk about pricing' resolves to a real timestamp.
- Silence and speech boundaries, so pauses, breaths and filler words can be cut on clean word edges instead of mid-syllable.
- A visual reading of the footage, so the subject can be kept in frame when the aspect ratio changes and cutaways can land on the right moment.
- The structure of the project itself: clip order, durations, tracks, and any effects already applied.
The layer an edit depends on decides how it behaves. Silence removal works from the transcript, which is why it is exact on a talking-head recording and a poor fit for a screen recording where the dead time is visual, not audible.
2. Turning One Sentence Into Several Edits
The part people notice is the translation step: one plain sentence becomes an ordered list of concrete operations. You describe the outcome, and the editor decides which operations produce it and in what order. Assembly comes first, then pacing, then motion and look, then text and audio, because an effect applied before a cut has to be redone after it.
- Cut the pauses and filler words on word boundaries.
- Switch the project to a 9:16 frame and keep the speaker centred.
- Generate word-timed captions, then apply a yellow caption look to all of them.
Ask for the result you want rather than the steps. Cubix picks the operations, runs them in a sensible order, and tells you what it changed.
3. Listening vs Watching
The single most useful thing to understand about AI editing is the difference between a request that language can answer and a request that only the picture can answer. 'Cut the bit where I say we raised a seed round' is a language question. 'Cut the bit where the page is loading' is a picture question, because nothing in the transcript marks it. Cubix routes these differently, and knowing which kind you are asking for makes your instructions land the first time.
- Language questions: quoted lines, topics, repeated sentences, false starts, pauses, filler words.
- Picture questions: on-screen waiting, spinners, clicking around, facial reactions, composition, whether the subject is cut off at the edge of frame.
- Mixed questions work too, e.g. 'zoom in when I mention the price' — the phrase locates the moment, the picture decides how the zoom should sit.
On screen recordings, demos and gameplay, describe what you can see rather than asking for silence removal. Most of that footage is unnarrated, so a purely audio-driven trim reads nearly all of it as removable.
4. Non-Destructive Timeline Changes
Every edit is written as instructions against your timeline — cut points, clip order, zoom and speed keyframes, caption timings, volume curves — not baked into your media. Your uploads are never rewritten, which is what makes it safe to try an aggressive pass and step back from it. Rendering only happens when you export, so the master files you uploaded stay exactly as they were.
- Trimming a clip that is already on the timeline moves its edges in place; it keeps its effects, transitions and identity.
- Splitting a clip leaves both halves editable, with any earlier speed, zoom or grade still attached.
- Passes you run from the quick edit rail record themselves on a card, so you can see what is applied and revert just that pass.
5. Reviewing, Adjusting, Undoing
Conversational editing is iterative by design, so the review step is part of the workflow rather than a fallback. Everything Cubix does in a single reply is grouped into one undo, so one Ctrl+Z (Cmd+Z) takes back the whole change rather than unpicking it one keyframe at a time. If the edit is close but not right, describing the correction is usually faster than undoing and starting over.
✖ Assuming an AI video editor generates synthetic footage of you
Why it fails: It confuses text-to-video generators, which invent new imagery, with an AI editor, which cuts and arranges the footage you actually recorded.
✔ Better approach: Cubix edits your real footage. It can also generate stock-style B-roll, images, music and voiceover as separate additions, but your recording stays your recording.
✖ Describing the tool instead of the outcome
Why it fails: Requests like 'use the blade tool at 01:12' assume a mechanical workflow and give the editor no idea what the cut is for.
✔ Better approach: Say what you want to be true of the finished video: 'cut from the end of the intro straight into the demo'.
✖ Asking for silence removal on a screen recording
Why it fails: The dead time in a demo is usually visual — loading, waiting, hunting for a file — and none of it appears in the transcript.
✔ Better approach: Describe it visually: 'speed up the parts where the page is loading and cut the bit where I am looking for the file'.
Execute this workflow automatically in Cubix
Upload a recording, describe the video you want, and watch the cuts, captions and framing land on the timeline.
Frequently Asked Questions
Common questions around this editing workflow.