Multi-track timeline
Video, audio, text and image tracks, trimming, splitting, stacking and frame-accurate navigation.
Import your own footage, then type or dictate the edit in ordinary language. The app works out which operations you mean and carries them out on your timeline, as one step you can undo with one press.
Clip Smasher does not generate footage. There is no prompt that produces video, images, scenes, avatars or voices. You import your own video, audio and images, and the natural-language controls carry out editing operations on that media: trimming, splitting, reordering, transitions, effects, overlays, speed changes and audio mixing.
What it replaces is the menu-hunting, not the shoot. “Put all the clips on the timeline and add a crossfade between every clip” is a long run of drags and taps expressed as one sentence, and the result is the same edit you would have built by hand.
The instruction is broken into clauses and each clause is matched against a grammar of editing operations that ships inside the app. A clause that matches becomes a concrete operation against a concrete clip. “The second one”, “the music” and “all of them” are resolved to actual items on your timeline before anything runs.
Because the grammar is in the binary rather than on a server, there is nothing to sign in to, nothing to wait for and no connection required. It also means the app is honest about its limits: it either knows a phrasing or it does not, and it never improvises an edit it is not sure you asked for.
Open the command panel from the header. Type in the box, or tap the microphone and say it. Dictation goes through your phone's own speech recognition and arrives as text you can correct before sending.
Every operation in the instruction is worked out and applied together against a copy of the project, and committed to your timeline only once the whole thing validates. The panel closes and the edit is on the timeline.
So the safety net is undo rather than a confirmation screen: one press takes the entire instruction back, however many separate edits it turned into.
Half an instruction you did not agree to is worse than none, so a clause it cannot place stops the whole command. It then asks about that clause specifically, with a box to retype only that part, and keeps everything it did understand.
A sample, grouped by what they do. Every command below is exercised against a real project in the app's own test suite, so these are not illustrations; they are the wording itself.
The vocabulary is much wider than the list. Speed, for example, answers to faster, slower, slo-mo, timelapse, hyperlapse, multipliers like 1.5x, and named factors like double and half. Text overlays answer to title, caption, subtitle, lower third, end card and countdown, in any of thirteen fonts, by name or by asking for “a handwritten font” and letting it choose.
A guess that lands on the wrong clip costs more time than a question, so several situations are questions on purpose:
There is also a small set of things it will not do because the editor itself cannot: playing a clip backwards, for instance, is not the same edit as reversing the order of the clips, so asking to “reverse it” asks which you meant rather than quietly rearranging your timeline.
The clause that stopped the command is written to a queue on the device and later uploaded so a rule can be written for it. File names are replaced with <file> before anything is stored, and nothing identifying you or your phone is attached. Identical wordings from different people are counted together, so the phrasings that defeat the most people get fixed first. What leaves the device, exactly.
No. Clip Smasher works with footage you import. Its natural-language controls perform editing operations on your media rather than generating new footage from prompts. It does not generate video, images, scenes, avatars or voices, and there is no prompt box that produces media out of a description.
Nothing is applied. The app identifies the specific part of the sentence it could not place, asks about that part on its own and gives you a box to retype just that clause. The rest of the instruction is kept, so you never have to retype a long command to fix three words.
Yes. Every operation in one instruction is applied together as a single step, and one press of undo takes all of it back. There is no half-applied state to unpick, because a plan either applies in full or does not apply at all.
No. The instruction is read on the phone. There is no account, no API key and no network call in the path between typing a command and seeing the plan, which is why it works in airplane mode. The one exception runs the other way: a clause the app could not place is later sent up, with file names removed, so a rule can be written for it.
Yes. Tap the microphone in the command panel and speak. Dictation uses your phone's own speech recognition to turn speech into text; Clip Smasher receives the text, not the audio, and keeps no recording.
Yes. An instruction can carry several clauses, such as “put all the clips on the timeline, chop out the parts where there is no talking, use pixelize transitions”. They are read in order and applied together. If any single clause cannot be placed, none of them are applied.
Video, audio, text and image tracks, trimming, splitting, stacking and frame-accurate navigation.
What Clip Smasher does with no connection, and precisely what does and does not leave your phone.
Every colour, blur, texture and transform effect in the app, and which of them keyframe.