Telegram: the file is compressed; lock picture before upscale
Upscale the keeper, not the miss.
Upscale the keeper, not the miss.
Upscaling a shot you will regen is paying twice. Lock the cut, then upscale the keepers.
add_veed_captions transcribes and burns styled subtitles. Run it on a miss, or before picture lock, and you pay for type on audio you will throw away.
generate_multi_speaker_speech builds a conversation in one call. Attaching it before picture lock orphans a timed dialogue on a recut.
generate_music plus attach_audio_to_video in mix mode scores a file. Scoring a miss, or scoring before picture lock, is a bed you will mute.
add_veed_captions times speech into burned-in subtitles. Doing that on a miss, or before picture lock, pays for a track you will recut.
add_video_captions burns one authored line for the whole file. Overlay it before picture lock and you typeset a miss.
add_timestamped_captions burns lines you wrote at start_sec and end_sec you set. Recut the video and every timestamp is a lie.
generate_speech plus attach_audio_to_video lays a script on a clip. Narrating a take you will regen, or narrating before picture lock, pays for a read you will orphan.
change_voice keeps timing and delivery, swaps the Cartesia speaker. Re-voicing a miss, or re-voicing before picture lock, pays for a performance on frames you will replace.
edit_video's aspect_ratio crops or letterboxes a file. Reformatting misses, or reframing before picture lock, spends editor_export on pixels you will regenerate.
edit_video can slow, speed, or ramp a described section. Ramping a miss, or ramping before picture lock, rewrites a clock you will throw away — and every later caption with it.
clone_voice_from_audio mints a Cartesia voice_id. Using that id on every experimental take, or cloning from a noisy scratch, wastes the clone and every generate_speech that follows.
cut_video stitches KEEP ranges to skip a middle section. Punching a miss, or punching before picture lock, spends a cut on a performance you will replace whole.
dub_video translates, clones delivery, and optionally lip-syncs. Run it before picture lock and every language version is a recut you will pay for again.
transcribe_audio (Cartesia ink-whisper) returns plain text. Running it on a throwaway soundtrack, or before picture lock, pays for a script of words you will never reuse.
generate_lipsync needs a face image and an audio file. Driving a rejected still, or a scratch read, spends a lipsync model on inputs you will replace.
attach_audio_to_video in replace mode mutes the original and lays a new track. Doing that on a miss, or before picture lock, destroys a soundtrack you might have needed and bills a swap you will repeat.
dub_video with translate_audio_only is a lighter language pass. It still bills against length, still runs async, and still dies if you recut the parent.
cut_video with one KEEP range is a clean trim. Trimming a miss, or trimming before picture lock, produces a short leftover you will not ship.
Define picture lock when a shot can still regenerate: a signed change-list, and which picture changes reopen grade and mix.