Guides

    Lipsync credits ignore the script, the voice and the language

    A 30-second talking head is rate times seconds; across the lipsync catalogue that rate spans a documented multiple, and trimming audio is the only way the bill falls.

    Versely Team6 min read

    A 30-second talking head is rate times seconds. Across the lipsync catalogue that rate spans a documented multiple, and trimming the audio is the only way the bill falls. The script, the face, the voice and the language change none of it.

    That formula is the entire 30-second talking head cost page. Treat that page as live. This post is the constraint the page is built on: the meter reads duration, not copy.

    Rate times seconds

    Lipsync on Versely bills as credits = rate_per_second × seconds. There is no entry fee, no minimum block, and no rounding up to the next five seconds on this shape. One second costs the rate. Sixty seconds cost sixty times the rate. Doubling the clip doubles the credits, exactly. That is not true of the block-priced or base-plus-marginal models used for generated video.

    seconds is the length of the finished clip, which for lipsync is the length of the audio you feed it. Trim the audio and the bill falls in exact proportion. A rambling 45-second read cut to a tight 30 removes exactly a third of the charge. Padding a 20-second line to fill a 30-second slot adds half again.

    rate_per_second is a property of the model, and on some models of the resolution band too. It does not move with word count, language, or which voice you picked. The AI lipsync tool is how you run those rows. The cost page is how you see the meter before you confirm.

    A worked example on the live page uses Sync Lipsync 2.0: 5 credits a second × 30 seconds = 150 credits. That is an anchor, not a recommendation. Cheaper and dearer rows use the same arithmetic.

    The documented multiple

    Every per-second (and resolution-banded per-second) lipsync rate in the catalogue sits on that same shape, and the rates are not close. At the time this was written the unit rate ran from 1 credit a second to 30 — a 30× spread on identical duration. A 30-second clip is therefore 30 credits or 900, depending only on which row you dispatched.

    The cost table is a rate card, not a ranking: one representative model per distinct price. Several models publish separate SD and HD/4K rates, so one family can sit at two places on the spread. Rows that do not bill per second are not on that card. Kling Lipsync is a per-5s token — do not read its headline as a per-second rate and multiply. The cost hub maps each billingType to the page that explains it. Quote the live scenario for the current model count.

    What the meter refuses to read

    Nothing in rate × seconds inspects the script. A 30-second Japanese read costs the same as a 30-second English read on the same model and band. Translating a finished video into other languages is a different job on a different formula.

    Nothing in it inspects the voice. Swapping a cloned take for a stock voice does not change the lipsync line. If you are generating the audio rather than recording it, that is billed separately, per 1,000 characters of script, on voiceover for a script. Two meters. The talking-head page does not include the speech charge.

    Nothing in it inspects the face. Uploading a still versus generating one is a different decision. The generated still, if you generate one, is an image job. The lipsync job still prices the seconds of mouth motion.

    Captions burned onto the finished clip are their own render with their own preset multiplier. They are listed on the talking-head page under what the charge does not cover, because they are not lipsync.

    The only way the bill falls

    Once the model and band are picked, duration is the only lever that lowers the charge. That is the claim. The other levers on the page either pick a different rate or raise the total:

    • Which model renders it. Same 30 seconds, 30 credits or 900. This is the largest swing, and it is a choice of rate, not a trim. Best lipsync model is the capability ranking over the same pool — input type, not price. It does not rewrite the formula.
    • Resolution, on models that price it. Fabric and Wan publish SD next to HD/4K. The gap between those two options is often as large as the gap between two different models. Picking SD is picking a lower rate. It is not trimming.
    • Retakes. Every retake is a fresh render at the full rate. There is no partial charge for a clip you throw away and no discount on the second attempt. A five-second test on the same row is the cheap insurance, because five seconds is five times the rate.

    Cut the audio. That is the only change that leaves the row and the band alone and still shrinks the bill, in exact proportion to the seconds you remove. Do not pay for silence you intend to crop later. A completed render is charged whether or not you like it. The formula does not have a quality term.

    FAQ

    Is a 60-second talking head exactly twice a 30-second one?

    Yes, on this meter. Per-second billing is linear with no entry fee and no minimum block, so doubling the duration doubles the credits precisely. That is not true of the block-priced or base-plus-marginal models used for generated video.

    Does the language of the script change the price?

    No. Nothing in this formula reads the language — only the duration. A 30-second clip costs the same in Japanese as in English. Translating a finished video is a different job on a different formula.

    Why is one lipsync model so much more than another?

    Because the rate is a property of the model, not of your clip. The catalogue publishes per-second lipsync rates from 1 to 30 credits, and every one of them is billed on the same rate-times-seconds formula. Cheaper is not better; the buyer's guides are where models get judged.

    Does generating the voice count in this formula?

    No. If you generate the audio, that is billed per 1,000 characters of script on the voiceover meter. Lipsync then bills the seconds of the finished clip. Two charges, two pages. The talking-head number is only the second one.