Script Timing Estimator in the browser is not a generate
/free-tools/script-timing-estimator never spends a credit. Do not pay a model to do an in-browser file job.
Script Timing Estimator never spends a credit. Paste a script, pick a words-per-minute pace, read the spoken duration and a suggested clip count. Typical narration presets are on the page. Nothing is sent anywhere. Do not pay a video model to "see how long this reads."
Generation on Versely always costs credits. There is no free generate tier. Arithmetic on your own text is not a generate.
Duration is division, not a render
Word count divided by a chosen pace is the whole product. The tool counts the words you paste, divides by WPM, multiplies by 60. That is the spoken-duration figure. It is not a performance. It does not know your pauses, your product names, or the breath before a line.
Presets on the page:
- Careful / tutorial — 120 wpm
- Conversational — 140
- Typical voiceover — 150 (the usual VO target)
- Commercial — 170
- Fast / hype — 190
The slider runs 80 to 220. 150 is a target, not a prediction of a specific speaker. If you already have a recording, time that rather than the script.
Paying Veo, a lipsync row, or a TTS model to discover that 420 words at 150 wpm is about 2 minutes 48 seconds is how you burn the expensive end of the stack on a sum. Run the estimator first. Then buy the voice.
The clip-count table is not a cut list
The second readout is duration divided by 8, 15, 30, or 60 seconds, rounded up. It does not know where a thought ends. It is a fit check against a format: will this narration cover four 15-second clips, or are you writing a 60-second VO for an 8-second hole.
That number is how you stop writing a paragraph for a Veo 4 / 6 / 8 second take. If the estimator says 40 seconds of speech, you do not have an 8-second problem. You have a script problem, or you need several shots. Cut on meaning after you know the script is in the right ballpark.
Do not prompt a cinematic model to "read this and time it." It will invent a scene. It will not give you a WPM table.
Credits start after the script fits
What it costs walks the real meters: per second of video, per 1,000 characters of TTS, per export. Those formulas belong to the take. The estimator is upstream of all of them. Fit the words to the slot while the cost is zero. Then pick the voice row, the picture row, the caption pass.
A practical order:
- Write the line you actually need on screen or in VO.
- Paste it here at 150 wpm, then at the pace you will really speak.
- If it overruns the slot, cut the script — not the model budget.
- Generate picture and voice only when the arithmetic says it fits.
Nothing about that sequence requires an account. Disconnect the network. The textarea still divides.
What this will not do
It will not transcribe a recording. It will not place captions. It will not know that "Q3" is one token and three syllables. It will not split on sentences for you — that is a different free tool if you need cues. Treat the duration as accurate for a steady read, and leave headroom for names and numbers.
FAQ
Is the estimator actually free, or is it a trial?
Actually free, with no account and no limit. It runs on the device. That is different from generating with an AI model on Versely, which costs credits.
Why is the duration different from my recording?
WPM is an average. Pauses, names, numbers, and a breath before a line all stretch real speech. 150 wpm is a typical voiceover target. Time the recording when you have one.
What does "suggested clip count" mean?
Duration divided by 8, 15, 30, or 60 seconds, rounded up. It is not an edit decision. If the table says six 15-second clips, you still cut on meaning. You just know the script is a six-clip script, not a one-clip script.
Should I estimate before I pay for TTS or video?
Yes. TTS is metered per 1,000 characters. Video is metered by what it renders. Both are wasted if the script cannot fit the slot. This page is a zero-credit check before those meters start.