Automate · Versely Agent

    Run a one-tap AI video template

    Attach a photo. Get the template's finished video.

    What you say to the agent

    No special syntax — just describe it like you would to a person.

    Run the finger snap template with this photo
    Turn this photo into the mugshot template video

    What it does, step by step

    1. 1

      The agent checks which inputs the template needs (usually a photo, attached via the chat's + button).

    2. 2

      Once every required input is gathered as a hosted URL, it starts the template run.

    3. 3

      A live-updating card tracks progress and shows the finished video once it's ready — no manual status polling needed.

    What it needs from you

    • Which template
    • The inputs the template needs (usually your photo)

    What comes back

    The finished video for that template, built from your uploaded photo(s).

    What it costs

    Each template lists its own estimated credit cost before you run it.

    Under the hood

    This is what the agent actually calls when you ask for it — real tools from its live surface, not marketing copy.

    start_template_runget_templateget_template_runcancel_template_run

    Kick off an AI Template run. Call this AFTER you have gathered every required input (per get_template / the template brief) as a hosted URL. For image inputs, ask the user to attach their photo via the chat composer's + button — the URL appears as attached_media in the user's next message and you can read it directly from the chat history. Do NOT call request_asset_upload for template inputs (that's only for workflow_assets table writes). On success, returns {run_id}; the frontend renders a TemplateRunCard that subscribes to realtime updates, so you do NOT need to poll get_template_run yourself in most cases.

    See it done in a real workflow

    Or start from a one-tap template

    Frequently asked questions

    What do I actually say to the agent to run a one-tap AI video template?+

    Just describe it in plain English — for example: "Run the finger snap template with this photo" The agent handles picking the right tool and model from there.

    What does the agent need from me first?+

    At minimum: Which template; The inputs the template needs (usually your photo). Anything else it needs, it asks for before running.

    What do I get back?+

    The finished video for that template, built from your uploaded photo(s).

    Does this cost credits?+

    Each template lists its own estimated credit cost before you run it.

    Is this one single action, or several?+

    Behind the scenes the agent may call more than one tool to pull this off — but you only ever describe the outcome you want in one message.

    You can also just ask for

    Ask your Versely agent to run a one-tap AI video template

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.