The brief is a sound, not a piece of music. A whoosh for a title, a footstep on gravel, rain on a window that can loop — the sound-effect job takes a description and optional duration and loop flags, and returns a short clip you place on a timeline.
Text-to-music writes a composition with genre and tempo; asked for cinematic hits it will still be a song, and you will hunt for a transient that was never spotted to your cut. Native-audio is soundtrack born with a video model, baked into the picture. Text-to-speech is a read. Mixing those jobs into one prompt is how you get a sung whoosh or a narrator describing rain.
Place the file and slip it a frame if the door is late. Regenerating the video so the bake accidentally contains the Foley you meant is the expensive diagnosis.
In practice
- Describe the event, the space and whether it should loop — 'whoosh' alone is a sample library guess.
- Generate the hit, drop it on the cut, then slip; do not re-roll the picture to chase a transient.
- Keep SFX off the speech job and off the song job; footsteps under a narrator are a second file.
The mistake to avoid
Prompting a music model for a whoosh, or regenerating the video so the native bake contains the hit you needed as a separate file.
Go deeper
A whoosh is generate_sound_effect, not a song prompt
Generate a sound effect is one agent job: a described hit, riser, or ambience. Using generate_music for SFX, or regenerating a video for one missing hit, wastes credits.
Related terms
Text-to-music
Text-to-music generates an original composition from a written description of genre, instrumentation, mood and tempo — with or without sung lyrics.
Native audio
Native audio means a video model generates its own soundtrack — dialogue, effects, ambience — in the same pass as the picture, rather than leaving you a silent clip.
Stem separation
Stem separation splits a finished mix into its component parts — vocals, drums, bass, other instruments — as separate audio files.
Text-to-speech
Text-to-speech converts written text into spoken audio using a synthetic voice you choose before generating.
Voice cloning
Voice cloning builds a reusable synthetic voice from a sample of a real one, so new scripts can be spoken in that voice later.
The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.