Linear per second

    What does a 30-second AI talking head cost?

    Between 30 and 600 credits. Lipsync bills per second of finished clip, and across the 17 lipsync models Versely runs the rate spans 1 to 20 credits a second — a 20× spread on the same 30 seconds. The script, the face, the voice and the language change none of it.

    Billed per second of finished clipper_second

    The formula

    credits = rate_per_second × seconds
    rate_per_second
    set entirely by the model, and on some models by the resolution band too — 1 to 20 credits across the catalogue. It does not move with script length, voice or language.
    seconds
    the length of the finished clip, which for lipsync is the length of the audio you feed it. Trim the audio and the bill falls in exact proportion.

    Worked example — A 30-second piece to camera on Sync Lipsync 2.0

    1. 1.rate_per_second = 5 credits
    2. 2.seconds = 30
    3. 3.5 × 30

    = 150 credits

    Why the same 30 seconds has a 20-fold price range

    Linear per-second billing is the easiest formula on Versely to predict and the easiest one to overpay on. There is no entry fee, no minimum block and no rounding: one second costs the rate, sixty seconds cost sixty times the rate. That makes the arithmetic trivial and makes the model choice the entire decision.

    The spread is what people miss. Every rate in the table below animates a face to speak, and every one is billed on the same rate-times-seconds shape — but the rates differ by 20×. On a single 30-second clip that is 30 credits against 600. Across five clips a week for a month it is 600 against 12,000.

    The same shape prices video extension and motion-control jobs, which is why a long extend can quietly become the largest single charge in an account: nothing in the formula caps it except the model's own maximum duration.

    The lipsync rate card, second by second

    Every figure below is computed from Versely's live price matrix — the same matrix the app quotes you from before it charges. These are prices, not verdicts: the buyer's guides are where models get ranked.

    Every per-second rate in the lipsync catalogue

    One representative model per distinct rate. A rate table, not a ranking — cheaper is not better, and the buyer's guides are where models get judged.

    Credits per secondExample model10s30s60s
    1VEED Avatar Audio103060
    4HeyGen Avatar V340120240
    5Sync Lipsync 2.050150300
    6Kling Avatar Standard60180360
    8VEED Fabric 1.0 (SD)80240480
    9Sync Lipsync 2.0 Pro90270540
    10Wan 2.2 Speech to Video (SD)100300600
    12Kling Avatar Pro120360720
    15VEED Fabric 1.0 (HD/4K)150450900
    17Sync React 11705101,020
    20Wan 2.2 Speech to Video (HD/4K)2006001,200

    Every cell is the rate multiplied by the duration — that is the entire formula. 22 rate rows across 17 models collapse into these 11 distinct prices.

    Where each model stops

    The catalogue publishes a floor and a ceiling per model. The ceiling is the rate times the longest clip it will render, so it doubles as a maximum-length check.

    ModelPublished credit rangeLongest clip that implies
    VEED Avatar Audio3 – 6060s at its top rate
    VEED Avatars3 – 7070s at its top rate
    VEED Lipsync4 – 8181s at its top rate
    LTX 2.3 Audio to Video50 – 20020s at its top rate
    Sync React 184 – 25115s at its top rate
    HeyGen Avatar V317 – 408102s at its top rate
    Sync Lipsync 2.025 – 600120s at its top rate
    Kling Avatar Standard29 – 675113s at its top rate
    Sync Lipsync 2.0 Pro42 – 1,000111s at its top rate
    HeyGen Avatar V550 – 1,200120s at its top rate
    HeyGen Image to Video50 – 1,200120s at its top rate
    Kling Avatar Pro58 – 1,380115s at its top rate
    VEED Fabric 1.040 – 1,800120s at its top rate
    VEED Fabric 1.0 Text40 – 1,800120s at its top rate
    Wan 2.2 Speech to Video50 – 2,400120s at its top rate
    Wan 2.2 Speech Turbo50 – 2,400120s at its top rate
    VEED Fabric 1.0 Fast50 – 2,400120s at its top rate

    What changes a lipsync bill

    Which model renders it

    The only lever that matters. Same 30 seconds, 30 credits or 600.

    Audio length

    Directly proportional. Cutting a rambling 45-second read to a tight 30 removes exactly a third of the bill.

    Resolution, on models that price it

    Several lipsync models publish separate SD and HD/4K rates, and the gap between them is often as large as the gap between two different models.

    Retakes

    Every retake is a fresh render at the full rate. There is no partial charge for a clip you throw away and no discount on the second attempt.

    What a lipsync charge does not cover

    • The voiceover itself — if you are generating the audio rather than recording it, that is billed separately per 1,000 characters of script.
    • Captions burned onto the finished clip, which are their own render with their own preset multiplier.
    • The face image, if you are generating that rather than uploading a photo.

    How many talking heads a month of credits buys

    A Monthly plan's 580 credits covers 19 thirty-second clips at the 1-credit rate, or 0 at the 20-credit rate. Same plan, same clip, 19 clips of difference.

    Plan sizes and what each grants are on the cost hub.

    Lipsync models on the per-second meter

    Every model page shows its own price matrix, its published credit range, and the option table this formula reads from.

    Frequently asked questions

    Is a 60-second talking head exactly twice the price of a 30-second one?+

    Yes. Per-second billing is linear with no entry fee and no minimum block, so doubling the duration doubles the credits precisely. That is not true of the block-priced or base-plus-marginal models used for generated video.

    Why is one lipsync model 20 times the price of another?+

    Because the rate is a property of the model, not of your clip. The catalogue publishes per-second lipsync rates from 1 to 20 credits, and every one of them is billed on the same rate-times-seconds formula.

    Does the language of the script change the price?+

    No. Nothing in this formula reads the language — only the duration. A 30-second clip costs the same in Japanese as in English. Translating a finished video into other languages is a different job on a different formula.

    What is the longest clip I can render?+

    Each model publishes its own ceiling. The second table converts every published maximum back into seconds at that model's top rate, so the length limits are readable at a glance.

    Do I pay for a render that comes out wrong?+

    A completed render is charged whether or not you like it. The cheapest insurance is a short test on the same model first — at 1 to 20 credits a second, a five-second test is a fraction of the full job.

    Different job, different arithmetic

    These are not variations on the formula above — each one meters something else entirely.

    Which model should I actually pick?

    This page prices the job. The buyer's guides rank the models, with the sort criterion stated on every one.

    Jobs billed this way

    Further reading

    See the price before you spend it

    Versely quotes the exact credit cost of every generation before it runs, and the agent will estimate a whole batch on request. The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.