How the clips get found, cut, and checked
Short answer
EasyClip watches and listens to the whole video in overlapping 20-minute sections, so it can use what’s seen and heard, not just what’s said. It compares every possible moment across the recording and keeps the best ones. Each cut is moved to the first word of its hook (or just before the action), cut from your original file, and checked again after export.
Updated
- 01
It watches and listens, not just reads a transcript
The actual video frames and the audio go to a multimodal AI model, one that takes in the picture and the sound together. So it sees the play, hears the reaction or the laughter, reads chat or a scoreboard when it’s legible on screen, and can pick moments from footage where nobody talks at all.
Every moment it finds is tagged by what carries it: something said, something seen, or something heard. That tag decides how the cut is placed and how the export is checked in the steps below.
- 02
It covers the whole video in overlapping sections
Long videos are split into 20-minute sections that overlap by 60 seconds, so a moment that straddles a boundary isn’t cut in half. Each section is watched and listened to separately, which means the third hour of a stream gets the same attention as the first ten minutes.
In each section, the model looks for moments a stranger could follow without context: a strong line, a reaction, a story with a payoff, a question and its answer, an attempt that lands or fails. If a section is genuinely dead, it returns fewer moments instead of padding.
- 03
It ranks every moment against the rest of the video
A moment isn’t kept just because it was found early. Candidates from every section are compared together, near-duplicates are removed, and the list is ranked best-first across the whole recording.
The top of that list becomes your included clips, about 10 per hour of video. Good moments below that line are listed as extras you can add afterward. If the video doesn’t have enough good moments, the unused slots are refunded instead of filled.
- 04
It moves each cut to the first word
The first pass only gives rough timestamps, and rough is how clips end up starting mid-sentence. So for every clip, the audio within 12 seconds of the rough start is searched for the first words of the hook. If they aren’t there, the search widens to 5 minutes either side.
The cut starts 0.15 seconds before that first word, so the line doesn’t sound clipped.
A moment carried by an action or a sound has no first word, so its start is placed on the nearest scene change or sound onset instead, a second or two before the payoff.
- 05
It cuts from your original file
Clips are cut from the file you uploaded, not from a compressed preview. Each one ends when the moment is over rather than at a fixed length. Every clip is then transcribed, and you get a transcript and an SRT caption file with it.
- 06
It checks the clip you’ll actually download
After export, the first 8 seconds of every clip that opens on a line are transcribed again and compared with the first 8 words of the hook it was supposed to start on. A clip that opens on an action or a sound has no words to match, so its first 4 seconds are watched and compared with the described opening instead. A timestamp can look right and still be wrong; the file can’t.
A clip that matches gets no note. The rest are marked “Double-check the first word” (or “the opening”) when they’re close, or “Needs a listen” (or “a look”) when they aren’t. Neither means the clip is bad, only that you should check the opening before you post it. When a spoken clip doesn’t match, it’s re-aligned once with a wider search before it’s delivered.
Not yet
What it doesn’t do yet
Vertical reframing
Clips keep the framing you uploaded. There’s no automatic 9:16 crop or facecam layout yet.
Burned-in captions
You get a transcript and an SRT file for each clip, but captions aren’t drawn onto the video.
Game-event detection
Nothing connects to kill-feed events, match data, or game APIs. Moments are picked from what’s seen and heard in the video.
Posting for you
It doesn’t schedule or publish to TikTok, YouTube, or Instagram.
Limits
Where it struggles
Speech is still the most reliable signal
Visual and sound moments are found, placed, and checked too, but a spoken first word is the easiest thing to cut on exactly. Commentary, conversation, and reactions give the most dependable results, and long silent stretches with nothing distinctive happening usually give fewer clips.
Overlapping voices
When several people talk at once, the first word is harder to pin down. Those clips are more likely to be marked check.
Music-heavy audio
Loud music under speech makes both finding and checking the opening less reliable.
Your footage
What happens to the video you upload?
The video is processed by an AI model that can watch and listen to footage (currently Google’s Gemini). The source file is kept for 7 days after the job finishes so you can add extras, then it’s deleted. Your finished clips and transcripts stay in your account until you ask us to delete them or close your account.
Nothing here promises views or followers. The goal is a batch of clips you’d actually post, with clean openings, so posting regularly doesn’t mean editing every night.
Questions about the method
Does it really watch the whole video?
Yes. Every 20-minute section of the video is reviewed, up to the upload limit of 12 hours and 8 GB, and the best moments are chosen across all of them.
Does it detect kills or other game events?
Not from game data. It isn't connected to kill feeds, match stats, or game APIs. It picks moments from what's on screen and what it hears, which is why the setup and the reaction usually make it into the clip.
Does it just read the transcript?
No. The model gets the video's frames and audio, so it can judge a play, a crash, a laugh, or a crowd, not only the words. Transcripts are used afterward: to find the first word of a spoken opening and to check the exported clip.
What do the notes on a clip mean?
They describe how closely the exported clip's opening matches what it was meant to start on: the first 8 seconds against the hook line, or the first 4 seconds against the described action or sound. No note is a close match. “Double-check” and “Needs a listen” or “Needs a look” are worth a quick check before posting.
Can it find moments with no talking?
Yes. A trick landing, a crash, a goal, a jumpscare, or a crowd roar can carry a clip on its own. Those clips start just before the action or sound, and their first 4 seconds are watched again after export. It's still a harder case than speech: footage where nothing distinctive happens for long stretches usually produces fewer clips, and unfilled slots are refunded.
How long is my video kept?
The source video is kept for 7 days after the job finishes so you can add extras, then it's deleted. Your finished clips and transcripts stay in your account until you ask us to delete them or close your account.
See it on your own footage.
Looking for your kind of video? Browse use cases.