What this does
The Cubix AI Agent combines speech understanding, visual understanding and timeline editing to turn a plain-English request into precise changes on your project.
Unlike video generation models, Cubix edits the footage you actually recorded. It will not invent you saying something you did not say.
How to do it
Follow these step-by-step instructions in Cubix.
It reads your footage first
When you upload, the agent builds a word-level transcript of what was said and a reading of what is on screen. That is what lets a request like 'the part where I talk about pricing' resolve to a real timestamp instead of a guess.
It works out which edits you meant
Your request becomes an ordered list of concrete operations. 'Make this vertical for TikTok with captions' becomes an aspect change, a crop that follows you, caption generation and a caption style — sequenced so that structural work happens before styling.
It edits your timeline and tells you what changed
The edits land on the same timeline you can work on by hand, and the agent reports what it did — how much it removed, where it cut, which style it applied — rather than just confirming it finished.
Listening vs watching
Questions that language can answer — a quoted line, a topic, a pause, a filler word — are answered from the transcript. Questions only the picture can answer — on-screen waiting, a reaction, whether you are cut off at the edge of frame — need the footage watched, which is a separate and slower step it takes when your request calls for it.
Pro Tips for Best Results
- It can see what you have already done. Passes you applied yourself are visible to the agent, so it will not blindly repeat them.
- You have the editor open while it works, so anything you change by hand is picked up on its next read.
- Its scope is video editing and Cubix help. It will decline general-purpose requests that have nothing to do with your video.
- When a request needs judgement it cannot make — which of two clips you meant, what order to merge things in — it asks rather than guessing, because guessing wrong on a destructive edit is expensive.
- It will never export on its own. Finishing an edit does not trigger a render; nothing leaves the editor until you ask.
The agent works on the same timeline you do, not a separate copy. That is why you can hand work back and forth freely: ask for a rough cut, adjust two clips yourself, then ask for captions, and nothing is lost in either direction.
Frequently Asked Questions
Common questions around this editing workflow.