Edit & Polish · Versely Agent

    Transcribe and caption my video

    Speech in. Styled, timed captions out.

    What you say to the agent

    No special syntax — just describe it like you would to a person.

    Add captions to this video in the "glass" style
    Caption this in Spanish with the "whisper" preset
    Preview the "fusion" caption style on the first few seconds before you caption the whole thing

    What it does, step by step

    1. 1

      Give the agent a video URL and a preset — it defaults to "glass" for a modern look if you don't specify a vibe.

    2. 2

      It transcribes the video's actual speech (the video must contain audible speech) and burns styled, word-timed captions using the VEED subtitle model.

    3. 3

      You can preview a style on a short clip first before committing to the full render, and it supports transcription in 165 language codes.

    What it needs from you

    • A video to work from
    • Which caption style/preset

    What comes back

    Your video with auto-transcribed, styled captions burned in.

    What it costs

    BASIC presets (21 styles) cost the standard rate; DYNAMIC presets (9 premium styles) cost double.

    Good to know

    • This transcribes real speech in the video. For a fixed caption you write yourself with no transcription, use the plain text-overlay capability instead.

    Under the hood

    This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.

    add_veed_captionspreview_caption_stylelist_caption_fonts

    Add styled, auto-transcribed captions to a video using the VEED Subtitles model (fal.ai veed/subtitles). Burns aesthetic preset styles (Glass, Whisper, Fusion, Glide, etc.) — different from add_dynamic_captions, which is the word-level animated style. Use when the user wants polished social-ready captions with a chosen visual look, or asks for VEED captions specifically. Also offer this proactively as a finishing step after any video / movie / workflow run that produces video content. Two preset tiers: BASIC (1× credit cost, 21 presets — simple/plain/beans/corpo/etc.) and DYNAMIC (2× credit cost, 9 premium presets — glass/whisper/glide2/fusion/glide/terminal/handwritten/backdrop/backdrop2). Default to preset 'glass' for a modern look unless the user specifies a vibe. Supports 165 language codes for transcription (en-US, en-GB, es-ES, es-MX, fr-FR, de-DE, it-IT, pt-BR, ja-JP, ko-KR, zh, ar-SA, ru-RU, etc.).

    The full tool behind it

    See it done in a real workflow

    Or start from a one-tap template

    Frequently asked questions

    What do I actually say to the agent to transcribe and caption my video?+

    Just describe it in plain English — for example: "Add captions to this video in the "glass" style" The agent handles picking the right tool and model from there.

    What does the agent need from me first?+

    At minimum: A video to work from; Which caption style/preset. Anything else it needs, it asks for before running.

    What do I get back?+

    Your video with auto-transcribed, styled captions burned in.

    Does this cost credits?+

    BASIC presets (21 styles) cost the standard rate; DYNAMIC presets (9 premium styles) cost double.

    Anything I should know before asking for this?+

    This transcribes real speech in the video. For a fixed caption you write yourself with no transcription, use the plain text-overlay capability instead.

    You can also just ask for

    Ask your Versely agent to transcribe and caption my video

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.