Time Captions to Shot Changes, Not Only Speech
Captions timed only to speech feel late at cuts and cover new faces. Shot-change rules, lead-in/out frames, and a QC pass that never covers a new speaker.
A caption that matches the waveform can still feel late. The cut happens, a new face fills the frame, and the previous speaker's last clause is still sitting on their mouth. Auto-timing from speech is doing the job it was built for — word-level timestamps, then chunks — and it has no opinion about the picture. Shot changes are a second clock. Readability is a different pass. This one is picture sync.
Speech timing and picture timing are different clocks
Speech timing answers "when was this said." Picture timing answers "when did the frame change, and is the speaker still in it." Section 6 of the BBC Subtitle Guidelines is explicit: it is less tiring if shot changes and subtitle changes occur at the same time, and many subtitles therefore start on the first frame of the shot and end on the last frame.
Netflix's public subtitle timing guidelines say the same thing in production language: subtitles should sit neatly within shots. They write the rules for 24 fps and then restate them in seconds so they travel: "half a second" is 12 frames at 24 fps, 15 frames at 30 fps. Versely's default video frame rate is 25 fps, so you should think in seconds and then snap to the nearest frame, not copy a 24 fps frame count blindly.
The failure on generated and social cuts is specific. A talking-head clip cuts to a product insert and the last clause is still on screen, so the product wears someone else's sentence. Or a new speaker starts on the cut and the old caption hangs over a mouth that is already forming different words.
The shot-change rules that actually hold
Two public guides already agree on the shape.
Prefer the cut as a caption boundary. If dialogue starts on the shot change, or within about half a second after it, Netflix wants the in-time on the first frame of the new shot. If an out-time is within half a second of the last frame before the cut, extend it to the shot change, leaving a two-frame gap. BBC's version is the same idea in editorial terms: match subtitles to the shot; if you must hang over a cut, do not snatch the caption off a few frames after it.
Do not straddle a cut for no reason. BBC: avoid a subtitle that starts in the middle of shot one and ends in the middle of shot two. Split the sentence at a linguistic break, or delay the next sentence to the cut. Netflix allows a caption to cross a shot change when the dialogue also crosses it. That is the exception, not the default. If the words end before the cut, the caption should end before the cut.
Do not carry a line into a new scene. BBC is blunt: never carry a subtitle into the next shot if that means crossing into another scene, or if it is obvious the speaker is no longer around. A reaction shot after a joke is the classic trap. The laugh is visual. Leaving the setup sentence on the reaction face both covers the face and spoils the cut.
Two-frame gaps, not 8-frame flickers. Netflix requires a minimum of 2 frames between subtitles at any frame rate, and at 24 fps closes gaps of 3–11 frames down to 2 frames. Gaps should be either 2 frames or at least half a second. BBC: if you leave a pause between two pieces of speech, make the gap at least a second, preferably a second and a half, or the result is jerky. At 25 fps, two frames is 80 ms; half a second is 12.5 frames. Snap to 12 or 13 and watch it back.
Lead-in, lead-out, and the two-frame gap
Lead-in is the in-time relative to the first audible frame. Netflix wants that in-time on the first frame of audio, or within 1–2 frames of it, unless a shot-change rule overrides. BBC wants appearance to coincide with speech onset, because lip-readers use the face as a cue. Do not let the caption arrive after the mouth has already started.
Lead-out is where speech-only timing is most often wrong. If there is no following subtitle, Netflix prefers the out-time about half a second past the end of the audio. If that extension would cross a shot change and the dialogue does not, prefer the shot. BBC allows up to about 1.5 seconds of hang after speech when the speaker is in shot, warns that a caption remaining too long gets re-read, and says a subtitle should not still be on screen after the speaker has disappeared.
When the two clocks conflict, picture usually wins at the cut, speech wins inside the shot.
- Snap in-times to speech onset, or to the first frame of the shot if speech starts on or just after the cut.
- Snap out-times to the last frame of the shot minus the two-frame gap, if the line would otherwise die just before the cut.
- Only then extend into a pause for reading time, and only if the speaker is still the subject of the frame.
- If reading time and the cut cannot both be satisfied, split the cue or cut words. Do not let a dense line ride across an unrelated shot so it "has enough time." That is the readability pass leaking into picture sync.
Never leave a caption over a new speaker
Speaker identity is a picture problem as much as a labelling problem. BBC: do not simultaneously caption different speakers if they are not speaking at the same time. New-speaker captions should come up as the new speaker starts. Netflix: do not reveal a punchline early where there is a visible reaction on screen.
On a generated dialogue scene the failure is mechanical. The model cuts on a new character while the caption engine is still flushing the previous clause. Anyone reading assigns the sentence to the face in frame.
The rule that survives a timeline:
- If the person in shot changes, the caption must change or clear on that frame, unless the same person is still speaking off-screen and you have labelled that.
- If two people alternate, do not let speaker A's out-time overlap speaker B's first frame.
- If the cut is a reaction — product, crowd, still face — clear the line before the cut even if that means a slightly fast out-time. BBC lists shot changes as a legitimate reason to give a subtitle less than the target reading time.
Forced narratives (on-screen type you are translating or repeating) follow the picture, not the voice. Netflix times those to the on-screen text itself, and if the text lasts the whole shot, the out-time is two frames before the shot change. Dialogue that fights that type is a placement problem: move the cue, or time to the type.
A QC pass on cuts, not on the waveform
Speech-timed auto-captions in Versely are a burn-in pass: add captions to a video transcribes, times, and composites. That is the right first pass, not a shot-change pass. Do it on a 25 fps timeline; a variable frame rate makes every snap-to-cut rule lie.
- Generate the speech-timed track. Use the caption tool, or the agent task to transcribe and caption.
- Preview the cuts. The editor's 480p preview pass is free and carries a short per-user cooldown. Scrub every hard cut for three defects: a line that starts mid-shot and dies mid-next-shot; a line on a new face; a line that hangs into a reaction shot.
- Split or snap. Where a cue straddles a cut, split at a clause boundary and put the second half on the first frame of the new shot. Where a cue dies 3–8 frames before a cut, extend it to two frames before the cut. Timed text overlays cover the cases auto-chunking will not: a punchline on the cut, a label that must not precede a bang, a silence caption.
- Watch once with sound off. If a caption still feels late, it is late relative to the picture.
- Export once. The editor is EDL-based: re-preview, then pay for a single final export regardless of clip count.
If the whole track is uniformly early or late after a trim, that is an offset, not a shot-change problem. The browser-side subtitle timing shifter moves every cue by a fixed amount on your machine and never calls a model. Progressive drift is a frame-rate mismatch; a constant offset will not fix it.
FAQ
Should a caption ever cross a cut?
Yes, when the spoken line itself crosses the cut and the same speaker remains the subject. Netflix allows that case. BBC would rather you split or delay. For brand and social cuts, treat crossing as the exception: different person, product, or scene change, end the cue first.
How many frames of lead-in should I add?
Netflix's tolerance is 1–2 frames before the first audible frame, unless a shot-change rule pulls the in-time to the cut. BBC wants coincidence with speech onset. More than about 1.5 seconds of anticipation, while the speaker is in shot, is too much in the BBC guide. Do not invent a house lead of several hundred milliseconds "for punch."
What if a shot is too short to read the line?
BBC's answer is to merge speech from two short shots and end the merged caption on the second shot change, unless merging spoils a gag. Netflix's is to borrow time by merging neighbouring cues, then re-segment. Neither guide tells you to let a fast cue ride across an unrelated shot just to hit a reading-speed cap. Cut words, or recut the picture.
Do burned-in social captions follow the same shot rules?
Yes. The rules are about what the eye does at a cut. A burned-in line that hangs over a new speaker is worse than a closed-caption one, because nobody can switch it off. Time it to the shot, then style it.