Guides

    Sora 2 Text to Video: write the sound into the prompt (4s, 8s, 12s, 720p, 10cr)

    Sora 2 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    Versely Team4 min read

    Sora 2 Text to Video has native audio. A silent prompt still gets a soundtrack you did not choose.

    OpenAI's row is text-to-video at 720p, 10 credits, durations 4s, 8s, 12s, 16s, and 20s, aspects 16:9 and 9:16. Catalog audio is true. There is no "picture only" toggle on the Versely card. If you describe a kitchen and say nothing about the room, you still get a room: hum, footsteps, a line of dialogue the model invented because the prompt left a hole.

    Write the sound, or the model will.

    The soundtrack is part of the sample

    Native audio means picture and track come from the same pass. On Sora 2 Text to Video that is not a premium add-on. It is the job. A prompt that only names blocking and light is an incomplete brief. The model will fill the mute with something that feels like a kitchen, and you will spend the next hour ducking a line nobody asked for.

    Name what we should hear, in the same sentence as what we should see:

    • Room tone: fridge, traffic, rain on the awning — or "near-silent studio, no speech."
    • Foley that belongs to the action: cup on wood, jacket zip, one footstep.
    • Speech only if the mouth is the shot, and then write the line.

    If you need a locked founder voice across a series, that is TTS after a silent-looking plate — and Sora is the wrong place to expect silence. The split is native audio versus TTS. Do not generate Sora, then stack a clone on a mouth that already spoke.

    Size the beat to 4, 8, or 12 seconds

    The useful windows for most social cuts are 4s, 8s, and 12s. 16s and 20s exist; they are not a place to hide a second scene. One continuous action, sized to the enum you picked. A three-act script in an 8-second clip is how the back half gets crushed into a smear.

    Camera language helps this family: who moves, which way, how the shot resolves. OpenAI's Sora tip wants scene narrative, camera movement, and temporal progression. Write "slow push-in, she sets the cup down, the room stays quiet" — not "cinematic masterpiece, epic vibe."

    Aspect is a control, not a vibe word. Set 16:9 or 9:16. Do not ask the prompt to "make it vertical."

    720p is the ceiling on this row

    Max output resolution is 720p. That is the product, not a draft waiting for a free 4K upscale inside the same generate. If the delivery spec is 4K theatrical, this is the wrong row, or you add a later upscale as a separate job. If the delivery spec is a 9:16 feed, 720p at 10 credits is the contract.

    The AI video generator is the launcher. The text-to-video ranking is the roster around it. OpenAI's family is Sora for motion and GPT Image for stills — do not ask Sora to lock a pack label the stills row should have locked first.

    A silent prompt is not cheaper. You still pay 10 credits. You just did not choose the track.

    FAQ

    Can I turn Sora 2's audio off?

    Not as a catalog control on this row. Audio is true. If you need a mute plate, say so in the prompt ("no speech, no music, quiet room") and still listen to the file. If you need a brand voice, generate with no spoken line in frame and add TTS, or pick a talking/lipsync row.

    Why do 4s, 8s, and 12s matter more than 16s and 20s?

    Because most cuts live there, and because a short enum rewards one beat. 16s and 20s are for a longer hold of the same action, not a second location.

    Is 720p a preview of a higher Sora tier?

    It is the listed max on Sora 2 Text to Video. Do not budget a 4K master from this card. If you need more pixels, that is a different model or a later upscale.

    Should I describe the sound in the same prompt as the picture?

    Yes. Native audio is one sample. Split briefs ("I'll add sound later") are how you get a soundtrack you did not choose, then fight it in the edit.