Timestamp blocks in Veo 3.1 prompts
Google's documented multi-shot method is timed prompt blocks, not a prose paragraph.
Write 0-2s / 2-5s / 5-8s. Do not write a prose paragraph and hope Veo cuts.
Google's documented multi-shot method for Veo is timed prompt blocks. Each block is a shot: camera, subject, action, sound. The blocks add up to the duration you actually requested — 4, 6, or 8 seconds on Veo 3.1. A paragraph that names three angles is not the same input. Veo will squash it into one wandering take, or it will invent cuts you did not budget, or it will spend six seconds on the first clause and two on the ending you cared about.
This is a grammar decision, not a vibe. Veo timestamp blocks are not Kling Omni shot lists and they are not Seedance choreography prose. The per-model surface is /prompting. Use this page to decide whether to timestamp, and how the blocks have to look if you do.
Why a paragraph fails on Veo's clock
Veo 3.1 generates 4, 6, or 8 seconds per call. That window is the whole movie. A paragraph with "then she turns, then we cut wide, then a crane up" is a three-act script for a model that will not pause to honour your "then." It has one denoise. Unscoped actions get compressed or dropped. The back half of your sentence is where endings go to die.
Timestamp blocks turn that window into a shot list the model is documented to follow. Google's own examples use ranges like [00:00-00:02], [00:02-00:04], [00:04-00:06], [00:06-00:08]. The production shorthand that survives a brief is 0-2s / 2-5s / 5-8s — slightly uneven on purpose when the land needs more time than the setup.
You are still in one generate. This is not three jobs. Identity continuity is the benefit. Per-beat still-lock is the thing you gave up. If beat two is a pack shot, stop. Timestamping will not make Veo honour a box it has not seen. Lock the still and cut.
The decision: timestamp, one shot, or don't use Veo
| You need | Write | Do not |
|---|---|---|
| One action, one camera, ≤8s | One shot. No blocks. | Fake timestamps for a single take |
| Two or three beats inside 8s, same person, cuts are the gag | 0-2s / 2-5s / 5-8s |
A paragraph with "then we cut" |
| Always-on dialogue and coverage | Blocks, with the line inside the block where the mouth is | Dialogue in a preamble the model applies to every second |
| A 15–30s continuous take | Wrong model or extend | A 8s timestamp prompt that pretends to be 30s |
| A pack / logo / approved still as a beat | I2V that still, then cut | A timestamp that "shows the product" |
| Kling-style six-shot storyboard | Omni, not Veo | Veo blocks as a fake storyboard |
Veo still wins always-on audio. Timestamping does not change that; it is how you aim the audio. Put SFX and the spoken line in the block they belong to, not in a header. The audio job is Veo 3.1 still wins always-on audio. The grammar job is this page.
How a block has to be built
Each range is a complete shot, not a residue of the last one.
0-2s: Medium shot, behind her, she pushes the vine aside.
SFX: leaves, distant birds.
2-5s: Reverse, face, the ruins behind her. She says, "This is it."
5-8s: Wide crane, she is small in the complex. No new dialogue.
Rules that hold up:
- The ranges must add up to the duration you set.
0-2 / 2-5 / 5-12on an 8-second generate is how the last beat gets eaten. If you need 12 seconds, you need extend or a second generate, not a longer last timestamp. - One camera instruction per block. "Medium, then a snap zoom, then handheld" inside
0-2sis a paragraph in a hat. - Dialogue lives in the block where we see the mouth. A line in a preamble is how you get speech over a wide shot of nobody talking.
- Do not re-describe the person in block three. Continuity is the point of one generate. Re-casting her in prose is how she recasts in pixels.
- Negative prompt is for exclusions (watermarks, extra limbs, text), not for "no cut in the middle of 2-5s." If you did not want a cut, do not timestamp.
Even blocks (0-2 / 2-4 / 4-6 / 6-8) are the default for four beats. Uneven blocks (0-2 / 2-5 / 5-8) when the land is the shot. Do not write six blocks into 8 seconds. Veo will smear them.
What timestamping is not
It is not extend. Extend is a continuation from a last frame. Timestamps are a plan for one unextended call. Do not timestamp an 8-second prompt as 0-8 / 8-15. The second range is not an extend request. It is a hallucination you typed.
It is not Movie Mode. A Veo file with three timed cuts is still one clip. A brand film with a scene plan, VO, and chained stills is a different object.
It is not a substitute for /prompting. Family tips still apply: Veo wants named camera moves, lighting, temporal flow. Put those inside the blocks. A timestamp with no shot language is just a clock.
The production rule
If the cuts are load-bearing — setup / reverse / land — timestamp. If there are no cuts, do not. If the cuts are load-bearing and one of them is a contractual still, do not timestamp that beat; still-lock it and edit. If you cannot fit the beats in 8 seconds without squash, you do not have a Veo timestamp job. You have a stitch job, an extend job, or a different model.
FAQ
Do I have to use [00:00-00:02] exactly as Google writes it?
No. Google's examples use that form; 0-2s is the same instruction in a brief. What matters is that each range is closed, sequential, and sums to the duration enum you actually sent (4s, 6s, or 8s). Fancy timecode will not add seconds the endpoint does not offer.
Can I timestamp a 4-second generate?
Yes, if you have two beats that fit. 0-2s / 2-4s. Three beats in 4 seconds is squash. Prefer one shot at 4s over three smeared ones.
Will timestamp blocks keep the same face across cuts?
Better than three separate generates, because it is still one identity event. Not as well as a reference still. If the face is the brand, condition the generate with a reference (or I2V a locked frame) and timestamp only the motion and the cuts. Do not ask timestamps to be a character sheet.
Should I copy this grammar onto Kling or Seedance?
No. Kling Omni wants a shot list in its storyboard fields. Seedance wants motion and choreography, and on the R2V endpoint it wants @Image1 tokens. Pasting 0-2s / 2-5s / 5-8s into another family is how you get a model that treats your clock as flavour text. Open /prompting for that model.