LTX 2.3 Retake Video: pick the row, not the brand (2s, 4s, 6s, 10cr)
LTX 2.3 Retake Video is one catalog generate. Use it because this job matches this row.
LTX 2.3 Retake Video is one catalog generate. Use it because this job matches this row.
LTX 2.3 Retake Video retakes a segment of an existing video. It takes a source video_url plus a prompt describing the desired change (the catalog example is "change flower to red rose"), plus optional start_time / duration / retake_mode (for example replace_audio_and_video). Category is edit-video. Ten credits is the catalog figure, billed per second. Durations are 2s, 4s, 6s, 8s, 10s, 12s, 14s, 16s, 18s, 20s. audio is null on the row flag. requires_image is false because the source is a video, not a still.
This is not "run LTX"
LTX is a brand with more than one slug. Best LTX model and the LTX provider roster are the family. This page is the retake row. A new scene from a paragraph is a generate. Continuing past the end of a clip is an extend. Redoing two seconds inside a clip you are keeping is a retake. Pick the row.
The prompt is an edit instruction, not a continuation. "Then she walks off screen" is an extend sentence. "Remove the man" / "change flower to red rose" is a retake sentence. Writing the former on this slug is how you spend 10 credits a second on a change the model was not asked to make in the right place.
The AI video generator will also list LTX. That does not make retake a text-to-video sample. You need a source clip. No source, no retake.
2s, 4s, 6s — even steps to 20s
The duration enum is even numbers from 2s through 20s. That is the segment you are replacing, not the length of the source file. A 2-second flower swap inside a 15-second plate is a 2s retake. A 20-second retake of a 20-second clip is a regenerate with extra steps. Start at 2s, 4s, or 6s unless the change actually spans longer.
retake_mode decides what the instruction is allowed to touch. replace_video keeps the original audio. replace_audio keeps the picture. replace_audio_and_video is the default in the description's example and will regenerate both if you leave it there. A visual-only note with the default mode is how the native (or recorded) stem you liked gets replaced by accident.
The row flag audio is null. Do not assume a soundtrack is created or preserved until you set the mode. If the plate has speech you need, replace_video is the mode that tries to leave it alone.
Ten credits a second on a segment you already paid for
The catalog credits value is 10, per second, with a matrix min of 20 and max of 200. A 2s retake and a 20s retake are different bills. That is the point of the even-step enum. Do not retake twenty seconds to change a flower in second four.
If there is no plate yet, you are on generate, not retake. If the plate is wrong as an idea — new room, new performance — retake cannot invent a camera that was never in the file. Generate. If the plate is right and one object is wrong, retake.
Video editing is still the cut, the title, the hard trim. Retake is new pixels inside a continuous segment. A hard cut is cheaper than a 10-credit-second rewrite of a line that should have been trimmed.
Use this row because the job is "change this object in this window of this clip." Do not use it because you like the LTX name.
FAQ
Can I retake from a still instead of a video?
This slug takes a source video and an edit prompt. requires_image is false. A still is a different category (edit-image or image-to-video). Retake is a segment of an existing clip.
Why even-numbered durations only?
That is supports_durations on this row: 2s through 20s in 2-second steps. There is no 5s option here. Size the segment to the enum or pick another model.
Will retake keep my original audio?
Only if you set retake_mode to replace_video. The catalog example names replace_audio_and_video as a mode. The default will not do you a favour. Set the mode to match the note.
What do 10 credits buy?
The catalog figure is 10, billed per second of the retake segment. You are buying an edit of a window inside a clip you already have, not a new text-to-video sample, and not an extend past the last frame.