Guides

    LTX 2.3 lipsync from a voice track is 20 credits, 1080p

    ltx-2-3-audio-to-video-lipsync is an audio-to-lipsync row at 1080p, 20 credits, native audio — it wants a voice track, not a premade avatar roster.

    Versely Team7 min read

    ltx-2-3-audio-to-video-lipsync is an audio-to-lipsync row at 1080p, 20 credits, native audio — it wants a voice track, not a premade avatar roster.

    A talking clip on this leftover starts with speech you already own. The mouth follows the WAV. You do not type the monologue into a text-to-video prompt and hope the model hears it. You do not pick a stock face from a library and call that a driving track. Make the audio. Then pick the row.

    A voice track is the brief

    Audio-to-lipsync means you already have speech as a file. Record it, clone it, or run a TTS row first. If you do not have a track yet, you are not on this leftover. Length usually follows the track. The catalog publishes an empty duration list because the WAV is the clock. Cut the line you will ship before you spend the floor on throat-clearing.

    AI lipsync is the launcher. Make a talking avatar video is the named agent phrasing. Cost is priced per model, shown before you confirm. Name LTX 2.3 Audio to Video if matching matters. A bare “make it talk” is how you land on a script-only twin or a premade roster.

    Bring a plate anyway when identity matters. requires_image is false on this slug — the job will run without a still — but a face you did not approve is a cousin mouth. A start frame is optional. Write performance around the uploaded track. Do not type the VO into the prompt and upload a different take.

    Hold the still (if you have one). Play the audio with your eyes closed. If either file would embarrass you, do not combine them. Changing the still later resets recognition. The audio is the performance. The model is not inventing the words.

    Health and finance stills do not get a photoreal talking presenter on this pass. The YouTube inauthentic-content bucket is the format, not the slug. Facility plates and overlay type stay available. A constructed clinician or advisor is not this leftover.

    LTX 2.3 is 20 credits, 1080p, native audio

    LTX 2.3 Audio to Video is the catalog row. Provider LTX. Content type lipsync. Category is audio-to-lipsync only. Catalog sticker is 20 credits, not discounted. Max output 1080p. Native audio is on — the track you supplied is the soundtrack. Aspects: auto, 16:9, 9:16. Start frame optional. Features listed include fast, high-quality, stylized, lipsync, audio-driven.

    That is the whole contract. Do not “upgrade in the prompt.” Change the row if you needed 4K, a premade cast, or a script with no WAV.

    6 published lipsync rows that take a driving audio track is the index this leftover sits on: Wan 2.2 Speech to Video, Kling Avatar Pro, LTX 2 Audio to Video, this LTX 2.3 slug, HeyGen Image to Video, and VEED Fabric 1.0. Bookmark those slugs. This page is only the 20-credit, 1080p LTX 2.3 card, and the roster it is not.

    Write the prompt around performance, not dialogue. The words are in the file. “Slight head nod, hold eyeline, no new room” is a shot. Pasting the script into the prompt while a different take is in the audio slot is how the mouth fights the mix.

    Caption anyway. Burned-in type is an editor job after picture lock. Auto-captions can transcribe this row because there is speech. Overlay is a different line — a CTA the mouth should not say. Do not ask LTX to typeset a headline into the wall behind the face.

    1080p is the ceiling. Do not treat this as a 4K master. If the still you bring is a sketch, the mouth will still follow the track, and the sketch will still be a sketch.

    VEED Avatars is a premade roster

    VEED Avatars is the other lipsync card in this brief, and it is not a driving-track row. Category is premade-avatars. Catalog sticker is 30 credits. Max output 4K. Native audio is on. requires_image is false. You pick a stock face from the product’s roster. You do not supply a WAV as the identity of the job. The library is the cast.

    That is a different door. Use it when you want a library presenter and you do not have a face you must keep. Do not treat LTX 2.3 as a cheaper VEED Avatars. LTX hears a file. VEED Avatars selects a body. Mixing the two in one brief is how a custom VO lands on a stock extra, or a roster pick waits for a WAV the row does not require as the brief.

    Sibling rows you should not confuse with this leftover:

    You have Published row Sticker
    Track first, 1080p, optional start frame LTX 2.3 Audio to Video 20cr, 1080p
    Track first, plate optional, flat 20 LTX 2 Audio to Video 20cr flat
    Stock face from a library VEED Avatars 30cr, 4K, premade
    Still + WAV, 720p talking still VEED Fabric 1.0 40cr, image required
    Photo plus a script, voice auto-generated VEED Fabric 1.0 Text 40cr, not “upload WAV”

    The three audio-to-lipsync rows without a /models page stay off this post. Do not invent their names. Live catalog is the published brief.

    Script-only is a different door

    A script-only talking-head row is text-to-lipsync or premade-avatars. You type words. The product invents or selects a voice. You did not supply a driving track.

    If you only have a paragraph, you are not in audio-to-lipsync yet. Make the track. Then pick LTX 2.3. Typing the monologue into Veo, Seedance, or a cinematic row and hoping the mouth appears is a generate of a scene, billed at that row’s rate, with glyphs that are atmosphere.

    Do not:

    • Paste the script into LTX 2.3 and skip the audio upload.
    • Pick VEED Avatars because the page said “lipsync” and you have a WAV you wanted followed.
    • Use a photoreal talking presenter on health or finance stills and label the face AI as if that left the bucket.
    • Treat 20 credits as a 4K master.

    Name the files, then the row. Track in. Optional start frame. 1080p. 20 credits. Native audio is the file you brought. A premade roster is a different SKU.

    FAQ

    Do I type the script into LTX 2.3 instead of uploading audio?

    No. The published category is audio-to-lipsync. It hears a file. Script-only is Fabric 1.0 Text, a premade twin, or another text-driven card. Make the track, then come back.

    How is this different from VEED Avatars?

    VEED Avatars is premade-avatars, 4K, 30 credits. You pick a library face. LTX 2.3 Audio to Video is audio-to-lipsync, 1080p, 20 credits, native audio from the WAV you supply. Roster versus track.

    Does LTX 2.3 require a still of the speaker?

    No. requires_image is false. Bring a start frame anyway when the face has to be someone you already approved. A missing plate is how you get a cousin mouth on a correct line.

    Is make-a-talking-avatar-video the same as this row?

    Make a talking avatar video is the in-app job for a face image plus audio. It still has to land on a row. Name LTX 2.3 if you want this 20-credit, 1080p, audio-to-lipsync card. A script with no WAV is a different job. A library presenter is VEED Avatars, not this leftover.