Guides

    SRT from Text in the browser is not a generate

    /free-tools/srt-from-text never spends a credit. Do not pay a model to do an in-browser file job.

    Versely Team4 min read

    /free-tools/srt-from-text never spends a credit. Do not pay a model to do an in-browser file job. Paste a script, set words per cue and a speaking pace, download a timed .srt. That is packing and arithmetic in the tab. It is not a video model, not a transcriber, and not a line on what it costs.

    Versely has no free generation tier. Every generate costs credits. This tool is free because it never calls a model. Prompting "turn this VO into subtitles" on a caption row is how a file job becomes a bill.

    The script already has the words

    The tool packs the spoken script into cues. Sentences become cue boundaries when they fit the words-per-cue cap; a long sentence is split rather than emitted as a 40-word caption. Each cue's duration is word count divided by the words-per-minute you set, with a 1-second floor so nothing flashes. Sequential SubRip indices, comma milliseconds, a blank line between cues. Positioning, ruby, and speaker tags are not generated — SubRip cannot carry them.

    Pace presets: 120 / 140 / 150 / 170 / 190 wpm. The slider runs 80–220. Words per cue is yours (1–30). None of those knobs invent a line. They clock the lines you pasted.

    Nothing is uploaded. No account. It keeps working with the network off. Same architecture as the rest of free tools: work on the device, no watermark. Treat the file as a starting track. Real speech is not even. Nudge after.

    Credits meter invention, not SubRip

    /cost groups credit formulas: per second, per thousand characters, per megapixel, per job, a caption preset multiplier. Script-to-SRT arithmetic is not one of them. Transcription is billed as captioning. A video generate that you hoped would paint subtitles is billed as video. The words were already in your doc.

    Add captions to a video is the billed sibling when the input is speech on a tape: 21 BASIC presets, 9 DYNAMIC at double, 165 language codes. That job transcribes. It does not typeset a script nobody said.

    If you already have a sidecar and only the format is wrong, use the SRT to VTT converter or the subtitle timing shifter. A constant offset is still a file job. Progressive drift is not a shift — that is frame-rate, and it needs the timings rescaled, not a model.

    Nudge the track; do not regenerate it

    The failure mode is "I need captions, so I will generate a video with the words in the shot." You will get glyphs, and you will have spent a credit formula on a job the tab already does.

    Use this page when you already like the words and only the clock is missing. Use a transcriber when the words live on the tape. Use a video model when you need frames. Players will accept standard SubRip. They will not accept it as finished if you never watched it against picture. Download script.srt, lay it on the cut, then shift.

    Do I already have the script, and I only need timestamps? Yes: pick a WPM, download, nudge. Clip and no script: transcription, quoted in credits before you confirm. Neither picture nor words: a generate, and captions are later. A missing clock is a word-count problem, not a prompting problem.

    FAQ

    Does SRT from Text spend credits if I am logged in?

    No. It never spends a credit. There is no file to upload and no model to meter. Logging into Versely elsewhere does not meter a script packer. Generation is the surface that shows a cost before you confirm.

    Why are the timestamps evenly paced?

    Each cue's duration is word count divided by the WPM you set, with a 1-second floor. Real speech is not even. Treat the file as a starting track to nudge, not as a finished caption file.

    Will players accept this SRT?

    It is standard SubRip: sequential indices, comma milliseconds, a blank line between cues. Positioning, ruby, and speaker tags are not generated, because SubRip cannot carry them. If you need WebVTT, convert after — still free, still in the browser.

    Should I generate captions with a video model instead?

    No. Scene models do not typeset a timed transcript. They sometimes paint illegible glyphs, and they bill as video. Use this page for a script you already have. Use billed captioning when the words are spoken on the tape.