Guides

    LTX 2.3 Audio to Video: the mouth pass after the picture exists (1080p, 20cr)

    LTX 2.3 Audio to Video wants a plate and a track. It is not how you invent the scene.

    Versely Team4 min read

    LTX 2.3 Audio to Video wants a plate and a track. It is not how you invent the scene.

    LTX, content type lipsync, category audio-to-lipsync, 20 credits, max output 1080p, audio: true. Description: "High-quality, fast AI video model for audio-to-video with lipsync, stylized transform capabilities." requires_image is false — the still is optional as a start frame, not the required identity lock some lipsync rows demand. What is required is the job shape: a track that already has the words, and a visual plate (prompt, and optionally a start frame) that already knows who is on camera. This SKU does not replace LTX 2.3 Image to Video Pro. It does not replace a text-to-video scene invent. It is the mouth pass.

    The track is the script. Do not type the script into the scene.

    Audio-to-video means the spoken words come from the uploaded track. Native audio is on. The visual prompt describes performance and framing around that track — gesture, setting, lighting — not the dialogue. Writing the VO into the prompt and uploading a different take is how the mouth fights the file.

    1080p is the listed max. Aspects: auto, 16:9, 9:16. No duration enum on the catalog record; length follows the track you fed it. 20 credits is the catalog figure (the matrix meters per second). A long VO is a long bill. Trim the track to the line you actually need before you buy the pass.

    AI lipsync is the door. Lipsync a photo or video to audio is the finishing route. Make a talking avatar video is the agent phrasing. LTX's roster is the family. Best lipsync model is the ranked shelf. The lipsync glossary is the mechanic.

    Optional start frame, required job discipline

    requires_image is false. You may pin a start frame. You may also prompt the visual. Either way, inventing the scene here is the expensive confusion. If you do not know who is on camera, generate the plate first — a stills row, or LTX 2.3 Image to Video Pro if you need motion without a mouth pass — then come back with a track.

    Stylized transform is in the description. That is a look you apply to a plate that exists, not a reason to skip the plate. Fast and high-quality are also in the description; they do not add a duration picker the catalog omitted.

    If the track is noisy, the mouth pass will be a noisy performance at 1080p. Clean the audio first. This is not a dialogue writer. This is not LTX 2.3 Retake Video. Retake edits a segment of an existing clip. Audio-to-video drives a talking picture from a track.

    After the picture exists

    Order of operations:

    1. Decide the plate. Optional start frame, or a prompt that is actually a person in a frame, not a plot.
    2. Decide the track. One line, trimmed.
    3. Run LTX 2.3 Audio to Video. 1080p, native audio, 20 credits as the catalog row.
    4. Caption in finishing. Lipsync is not typesetting.

    If step 1 is "we'll see what the model invents," you are on a generate SKU and calling it lipsync. The mouth pass will then be a mouth on a scene nobody approved. 20 credits still posts.

    FAQ

    Does this model require an image?

    No. requires_image is false. A start frame is optional. A driving audio track is the load-bearing input. If you have no track, you are not on this row.

    What resolution do I get?

    Max output is listed at 1080p. Aspects are auto, 16:9, and 9:16. Do not brief 4K from this SKU.

    Is this how I write a scene from a sentence?

    No. That is text-to-video. This row is audio-to-lipsync with native audio. Invent the plate elsewhere. Put the mouth on it here.

    How long can the clip be?

    The catalog does not list a duration enum. Treat length as driven by the track, and remember 20 credits is the listed row while the matrix meters per second. Shorten the VO first. Do not upload a podcast and hope the SKU becomes an episode generator.