Linear per second

    What does a 30-second AI talking head cost?

    Between 30 and 720 credits. Lipsync bills per second of finished clip, and across the 21 lipsync models Versely runs the rate spans 1 to 24 credits a second — a 24× spread on the same 30 seconds. The script, the face, the voice and the language change none of it.

    Billed per second of finished clipper_second

    The formula

    credits = rate_per_second × seconds
    rate_per_second
    set entirely by the model, and on some models by the resolution band too — 1 to 24 credits across the catalogue. It does not move with script length, voice or language.
    seconds
    the length of the finished clip, which for lipsync is the length of the audio you feed it. Trim the audio and the bill falls in exact proportion.

    Worked example — A 30-second piece to camera on Sync Lipsync 2.0

    1. 1.rate_per_second = 4 credits
    2. 2.seconds = 30
    3. 3.4 × 30

    = 120 credits

    Why the same 30 seconds has a 20-fold price range

    Linear per-second billing is the easiest formula on Versely to predict and the easiest one to overpay on. There is no entry fee, no minimum block and no rounding: one second costs the rate, sixty seconds cost sixty times the rate. That makes the arithmetic trivial and makes the model choice the entire decision.

    The spread is what people miss. Every rate in the table below animates a face to speak, and every one is billed on the same rate-times-seconds shape — but the rates differ by 24×. On a single 30-second clip that is 30 credits against 720. Across five clips a week for a month it is 600 against 14,400.

    The same shape prices video extension and motion-control jobs, which is why a long extend can quietly become the largest single charge in an account: nothing in the formula caps it except the model's own maximum duration.

    The lipsync rate card, second by second

    Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.

    Every per-second rate in the lipsync catalogue

    One representative model per distinct rate. A rate table, not a ranking — cheaper is not better, and the buyer's guides are where models get judged.

    Credits per secondExample model10s30s60s
    1VEED Avatar Audio103060
    3HeyGen Avatar V33090180
    4Sync Lipsync 2.040120240
    5Kling Avatar Standard50150300
    7Sync Lipsync 2.0 Pro70210420
    8Wan 2.2 Speech to Video (SD)80240480
    10Kling Avatar Pro100300600
    11LTX 2.5 Audio to Video Fast110330660
    12VEED Fabric 1.0 (HD/4K)120360720
    14LTX 2.5 Audio to Video Pro140420840
    16Wan 2.2 Speech to Video (HD/4K)160480960
    24Avatar X Reference to Video2407201,440

    Every cell is the rate multiplied by the duration — that is the entire formula. 26 rate rows across 21 models collapse into these 12 distinct prices.

    Where each model stops

    The catalogue publishes a floor and a ceiling per model. The ceiling is the rate times the longest clip it will render, so it doubles as a maximum-length check.

    ModelPublished credit rangeLongest clip that implies
    VEED Avatar Audio2 – 4848s at its top rate
    VEED Avatars3 – 5656s at its top rate
    VEED Lipsync3 – 6565s at its top rate
    LTX 2.5 Audio to Video Pro68 – 13610s at its top rate
    LTX 2.3 Audio to Video40 – 16020s at its top rate
    Sync React 167 – 20114s at its top rate
    LTX 2.5 Audio to Video Fast52 – 20819s at its top rate
    HeyGen Avatar V314 – 327109s at its top rate
    Sync Lipsync 2.020 – 480120s at its top rate
    Kling Avatar Standard23 – 540108s at its top rate
    Sync Lipsync 2.0 Pro34 – 800114s at its top rate
    HeyGen Avatar V540 – 960120s at its top rate
    HeyGen Image to Video40 – 960120s at its top rate
    Kling Avatar Pro46 – 1,104110s at its top rate
    VEED Fabric 1.032 – 1,440120s at its top rate
    VEED Fabric 1.0 Text32 – 1,440120s at its top rate
    Wan 2.2 Speech to Video40 – 1,920120s at its top rate
    Wan 2.2 Speech Turbo40 – 1,920120s at its top rate
    VEED Fabric 1.0 Fast40 – 1,920120s at its top rate
    Avatar X Reference to Video120 – 2,880120s at its top rate
    Avatar X Text to Video120 – 2,880120s at its top rate

    What changes a lipsync bill

    Which model renders it

    The only lever that matters. Same 30 seconds, 30 credits or 720.

    Audio length

    Directly proportional. Cutting a rambling 45-second read to a tight 30 removes exactly a third of the bill.

    Resolution, on models that price it

    Several lipsync models publish separate SD and HD/4K rates, and the gap between them is often as large as the gap between two different models.

    Retakes

    Every retake is a fresh render at the full rate. There is no partial charge for a clip you throw away and no discount on the second attempt.

    What a lipsync charge does not cover

    • The voiceover itself — if you are generating the audio rather than recording it, that is billed separately per 1,000 characters of script.
    • Captions burned onto the finished clip, which are their own render with their own preset multiplier.
    • The face image, if you are generating that rather than uploading a photo.

    How many talking heads a month of credits buys

    A Monthly plan's 774 credits covers 25 thirty-second clips at the 1-credit rate, or 1 at the 24-credit rate. Same plan, same clip, 24 clips of difference.

    Plan sizes and what each grants are on the cost hub.

    Lipsync models on the per-second meter

    Every model page shows its own price matrix, its published credit range, and the option table this formula reads from.

    Frequently asked questions

    Is a 60-second talking head exactly twice the price of a 30-second one?+

    Yes. Per-second billing is linear with no entry fee and no minimum block, so doubling the duration doubles the credits precisely. That is not true of the block-priced or base-plus-marginal models used for generated video.

    Why is one lipsync model 24 times the price of another?+

    Because the rate is a property of the model, not of your clip. The catalogue publishes per-second lipsync rates from 1 to 24 credits, and every one of them is billed on the same rate-times-seconds formula.

    Does the language of the script change the price?+

    No. Nothing in this formula reads the language — only the duration. A 30-second clip costs the same in Japanese as in English. Translating a finished video into other languages is a different job on a different formula.

    What is the longest clip I can render?+

    Each model publishes its own ceiling. The second table converts every published maximum back into seconds at that model's top rate, so the length limits are readable at a glance.

    Do I pay for a render that comes out wrong?+

    A completed render is charged whether or not you like it. The cheapest insurance is a short test on the same model first — at 1 to 24 credits a second, a five-second test is a fraction of the full job.

    Different job, different arithmetic

    These are not variations on the formula above — each one meters something else entirely.

    Which model should I actually pick?

    This page prices the job. The buyer's guides rank the models, with the sort criterion stated on every one.

    Jobs billed this way

    Further reading

    See the price before you spend it

    Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync — in your browser or on your phone.