Four hours of master. Forty cuts due by morning.
Exemplary AI reads the master once, picks the moments your rules allow, holds the speaker in frame at every aspect ratio, and captions them on brand in 120+ languages. Drafts land in a review queue, and nothing reaches a destination until a person signs it off.
The archive is full. The output queue is not moving.
Nobody in this market is short of material. What runs out is the number of hours a small team can spend turning material into finished, on-brand, on-policy cuts.
Nobody can reach the archive
Years of masters catalogued by filename, date and a one-line slug. If nobody remembers a moment, it is not in there. The rights were paid for either way.
One desk serves every title
A handful of editors cover every show, every fixture and every campaign. The queue gets triaged by whoever asks loudest, and the long tail never gets cut at all.
Every platform wants its own shape
Vertical here, square there, wide on the site, and a different length for each. One moment becomes four exports, reframed and recaptioned by hand every time.
Brand discipline breaks at volume
A clip a week can be checked by eye. Two hundred cannot. Type drifts, safe areas get ignored, and the logo ends up underneath the platform's own controls.
The rules live in people's heads
What may be clipped, what must never be, whose face needs consent on file, which sponsor cannot sit beside which story. Agreed once, then applied from memory at speed.
Localisation is priced per pass
Each language is commissioned separately from the same master, weeks apart. The versions start drifting from one another on the first day.
A ten out of ten that breaks a hard rule is excluded, not down-ranked.
Write your editorial and legal code once, in plain language. It compiles into named moment types and the standing rules a clip may never break. Other clippers rank moments. This one obeys you, and shows its working.
Candidates, with the reason attached
- 01:12:38Turning pointscore 9.1Selected
That is the moment it turned.
Inside the band at 34s. Scene boundary at both ends.
- 00:24:06Quotable linescore 9.8Excluded
Somebody knew, and said nothing.
Rule 2. The response that followed is not in the cut.
Highest score in the set, and excluded all the same.
- 02:03:55Crowd momentscore 7.4Selected
The whole bench is on its feet.
Crowd audio peaks across a scene change. No dialogue in the shot, and no restriction on the asset.
- 00:51:12Explainerscore 8.6Needs a person
Here is why the rule changed.
Rule 3. A minor is on camera and no consent is recorded.
Every candidate comes back with its type, its hook, its score and the rule that decided it, and the batch renders the selected ones as drafts.
One selection brain at two speeds. The same logic runs in the overnight batch and in the interactive pass, so "cut the whole weekend" and "find me the moment around 12:30" cannot come back disagreeing with each other.
It renders a frame, looks at it, and fixes what it got wrong.
Making the clip is the easy half. The agent renders frames from the finished cut, scores them with a vision model, repairs what fails and renders again. Where it could not check something visually, it says so.
- Legible at watch sizepasspass
- Contrast against the picturepasspass
- Clear of the control bandfailpass
- Subject not occludedpasspass
- Type and colour on brandpasspass
The caption sat inside the control band. Raised it, rendered again, checked again.
Verified on the delivered file
The same pass runs on the rendered file, not only on this preview. A clip that could not be verified visually is labelled as unverified rather than left to look fine. Either way it lands in the review queue as a draft.
Scored against a rubric, not a feeling
Legibility at the size it will actually be watched, contrast against the picture behind it, occlusion of the subject, and whether the type and the colours match your kit.
The platform's own furniture counts
Feeds paint their controls over the bottom of the frame. A caption that lands underneath them has failed the check, however correct it looked in the editor.
Checked on the delivered file
The same pass runs on the rendered output, not only on the preview, and the result rides on the clip as a badge. Anything that was not visually verified is labelled as not verified.
A seated speaker who looks down disappears from a face detector.
That is the failure every automatic crop makes sooner or later, and why reframing here does not run on vision alone. Face and body tracking are fused with the transcript's speaker turns, so the crop follows whoever is actually talking.
Illustrative frames. Every clip carries this report: what was detected, against what was rendered.
Vision under-counts. Audio over-counts.
Look down and the detector loses you: vision alone sees one speaker where there are two. Audio alone hears three, because a caller has no seat on screen. The count is the visual tracks raised by the audio, and only as far as there are real anchors in the picture.
A locked crop, or a pan with a clamp
A subject who stays put gets a fixed crop rather than a camera that breathes. A subject who moves gets a pan with a speed limit, because nothing gives away an automatic crop faster than one that snaps.
It reports what it saw against what it rendered
Two speakers detected, a split rendered. One detected, the frame held. When it gets a shot wrong you can see that it got it wrong, which is the difference between a tool and a black box.
The preview, the editor and the rendered pixels run the same code.
21 built-in styles, 8 word animations, and an AI style designer that lints its output before it ships: a caption needs a backing box, a shadow or a stroke. Emphasis tagging recolours the words that carry the line, timings untouched.
A style the designer cannot make legible does not ship. It comes back with the reason it failed.
Caption tracks are produced per language, with exactly one burned into the picture and the rest travelling as files. Brand kits and layout templates decide the type, the colours and where the logo sits, per output rather than per editor.
Five hundred files overnight, quoted before you commit.
Five hundred files is not five hundred jobs run one behind another. One request becomes one task per clip, each retrying on its own, so a corrupt file at 3am costs you that clip, not the night.
Separately tuned queues
- Deadline workshort clips, front of the line
- Long exportshour-long renders, their own lane
Before you submit
- Tasks
- 43
- Estimated time
- 38 min
- Estimated cost
- 43 task credits
Reframing and captioning never wait on each other, so the slowest single step sets the clock rather than the sum of all of them. Credits price per contract, and the estimate uses the same arithmetic as your invoice.
The overnight run and the one-off question agree
Ask about a single timecode at 11am and you get the same judgement the 2am batch made, because it is the same rulebook running at a different size.
Queues that do not block each other
Hour-long exports and deadline work sit in separately tuned queues, so the long job never parks itself in front of the urgent one.
Your archive is not a bucket
Read and write your own S3, GCS or Azure storage on both ends, and connect to the systems the archive lives in: your MAM or DAM, your watch folders, and your NLE, so a finished clip comes back as something an editor can open, not a file in a folder.
Live is a different job, and we will say so
Today the unit of work is a completed or growing file, not a live feed. If you need a goal cut, captioned and out inside ninety seconds while the match is still running, ask us on the call and we will tell you plainly where we are.
Forty drafts are waiting at 07:00. Somebody has to own them.
The overnight run does not deliver clips. It fills a queue, split by destination, with the rule that decided each moment still attached to the draft that came out of it.
Sports social, oldest first
- 02:03:55Crowd moment9:16Approved
Signed off by the sports desk at 07:04, and delivered to the destination it was queued for.
- 01:12:38Turning point1:1Sent back with a note
Start it on the whistle, not two seconds after. Re-cut against the note and back in the queue as a new draft.
- 00:51:12Explainer9:16Waiting on a person
Rule 3. A minor is on camera and no consent is recorded against the asset.
Nothing here reaches a destination until a row is signed off. The rule that decided the moment travels with the draft, so whoever approves it can see what the agent was obeying.
The queue knows whose desk it is
Drafts route by destination, language and library, so the sports desk opens sports and the duty editor opens the rest. Groups decide who can see a queue at all, which is the same control that decides who can reach the archive behind it.
A rejection travels with a note
Send one back with the reason: start it on the whistle, lose the first two seconds, that caption breaks the style guide. The note stays on the clip, the agent re-cuts against it, and the new draft arrives in the same thread, not as a fresh file nobody can trace.
Approve forty in one pass
Filter to a channel, work down the list, and approve what is right in one action. Every approval, rejection and note is recorded against the person and the clip, so the trail reads as a record of decisions rather than a record of deliveries.
Every version of that minute is indexed to the same frame.
Transcription indexes the speech; scene boundaries, audio energy, crowd reaction and burned-in graphics index the rest. Clips, caption tracks and translations point at the same frames, so versions cannot drift. Ask for the moment, get the timecode, not a filename.
Timecoded, attributable text, in 120+ languages.
Speaker diarization with word-level timecodes, so a search result is already a clip in-point and a translated caption still lands on the syllable. Speaker names are proposed only where the audio evidence supports them, and left as roles when it does not.
- Speaker diarization with word-level timecodes
- Translation on caption-grade timing across 120+ languages
- Filler words detected and removed on request
- Host01:12:31You said at the start of the year that the number would hold.
- Guest01:12:38And that is the moment it turned. Nobody expected the second revision.
- Host01:12:46We will come back to that after the break.
Clips, caption tracks and the copy around them are all cut from this one pass.
Scene detection
Cuts land on a boundary rather than mid-scene, so a clip does not open on half a shot.
Stories
Several moments stitched into one reel, in the order you want them read.
Exports
SRT, VTT, TXT, JSON, PDF and DOCX, alongside the rendered clips themselves.
Masters and unreleased material never leave the facility.
Rights, embargoes and unreleased footage are commercial assets before they are files. Uploading them to somebody else in order to make them searchable is the wrong trade.
The model comes to the storage
Rushes and masters stay put; inference runs beside them on models you host, including the frame scorer and caption designer, so no frame or transcript reaches a third party. Routing a step to a hosted model is a setting you turn on deliberately, and every call that leaves is logged.
Embargoes and rights hold
Groups scope which team and which agent reach which library. Restricted material stays invisible to everyone it should be invisible to, and access is revocable in a click.
The sign-off is a record, not a habit
Every approval is written down: who signed a draft off, what they sent back, and what they wrote on it. Months later, when somebody asks who cleared a cut, the answer is a row rather than a memory.
Deployed inside your own facility: on-premise, private cloud, or fully air-gapped. SOC 2 Type II, ISO 27001 and GDPR-ready.
Everything the platform does with a video file.
One deployment. The same machinery that reads your documents cuts, frames, captions and renders your video.
- Ingest
- Watch folders, upload, API, and MAM or DAM connectors
- Speech
- Transcription, speaker diarization, word-level timecodes
- Signals
- Scene boundaries, audio energy, crowd reaction, burned-in graphics
- Identity
- Speaker names proposed only where the audio supports them
- Language
- 120+ languages, translated on caption-grade timing
- Selection
- Named moment types, length bands and standing rules
- Editing
- Filler-word removal and scene detection for clean cuts
- Framing
- Auto-reframe to 9:16, 1:1, 4:5 and 16:9 with speaker tracking
- Captions
- 21 styles, 8 word animations, per-language tracks
- Emphasis
- AI tagging that recolours words without touching timings
- Brand
- Brand kits, layout templates, audiograms in four ratios
- Assembly
- Stories that stitch several moments into one reel
- Verification
- Vision-scored frames on the delivered file, not the preview
- Sign-off
- A review queue per destination, with every decision recorded
- Jobs
- Task-graph batches with cost and time estimates up front
- Storage
- Your own S3, GCS or Azure buckets on both ends
- Round trip
- Clips return as assets your MAM and your editors can open
- Exports
- SRT, VTT, TXT, JSON, PDF and DOCX
- Retrieval
- Ask the archive and get the segment, the speaker and the timecode
- Access
- Groups and IAM, scoped per team or library
- Systems
- 170+ integrations via MCP, plus custom MCP servers
- Deployment
- On-premise, private cloud or air-gapped
What content operations teams ask first.
Bring four hours and your style guide.
Send us one master and the editorial rules you already work to. We will compile them, run the selection, and show you the candidate list with the rule that decided each moment, including the ones it refused to cut. Judge it on your own material, not ours.
- Your rulebook, compiled and read back to you in plain language.
- Every candidate with its type, its length, its score and the rule that decided it.
- The moments it refused to cut, and which rule closed each one.
Run it inside your own facility, or send a single master under whatever paperwork your legal team wants first.