Speech, voice and audio

    Text-to-SFX

    Also called Sound-effect generation, Text to sound effect.

    Text-to-SFX generates a short sound effect — a whoosh, hit, footstep or ambience — from a written description, rather than composing a song.

    The brief is a sound, not a piece of music. A whoosh for a title, a footstep on gravel, rain on a window that can loop — the sound-effect job takes a description and optional duration and loop flags, and returns a short clip you place on a timeline.

    Text-to-music writes a composition with genre and tempo; asked for cinematic hits it will still be a song, and you will hunt for a transient that was never spotted to your cut. Native-audio is soundtrack born with a video model, baked into the picture. Text-to-speech is a read. Mixing those jobs into one prompt is how you get a sung whoosh or a narrator describing rain.

    Place the file and slip it a frame if the door is late. Regenerating the video so the bake accidentally contains the Foley you meant is the expensive diagnosis.

    In practice

    • Describe the event, the space and whether it should loop — 'whoosh' alone is a sample library guess.
    • Generate the hit, drop it on the cut, then slip; do not re-roll the picture to chase a transient.
    • Keep SFX off the speech job and off the song job; footsteps under a narrator are a second file.

    The mistake to avoid

    Prompting a music model for a whoosh, or regenerating the video so the native bake contains the hit you needed as a separate file.

    Go deeper

    A whoosh is generate_sound_effect, not a song prompt

    Generate a sound effect is one agent job: a described hit, riser, or ambience. Using generate_music for SFX, or regenerating a video for one missing hit, wastes credits.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.