A whoosh is generate_sound_effect, not a song prompt
Generate a sound effect is one agent job: a described hit, riser, or ambience. Using generate_music for SFX, or regenerating a video for one missing hit, wastes credits.
Describe the sound. Get the SFX. Generate a sound effect is one named agent job: generate_sound_effect from a prompt, optionally with duration and loop. A deep cinematic whoosh for a title, looping rain on a window, a UI click. That is the list. It is not a song, not a voiceover, not a reason to regen the whole clip because the door closed one frame early.
The tool requires a prompt. Optional: model, loop, duration, prompt influence, output format. You get a short audio clip of the described sound, ready to drop on a timeline. Related music tooling lives on the AI music generator page; that catalog is not this job.
Hits are not beds
generate_music writes a track. Asked for "cinematic hits in the score," it will still be a piece of music. You will then hunt for a transient that was never spotted to your cut. Generate the whoosh. Place it. Slip it a frame if the door is late. Regenerating the video so the picture accidentally matches a baked-in ambience is the expensive diagnosis.
Same split against speech. A narrator is generate_speech. Footsteps under the narrator are SFX. Mixing those two into one "make the audio" prompt is how you get a sung whoosh or a spoken rain loop.
Native-audio video models already bake some room. Use this job for the effect the bake missed: the riser on the cut, the specific impact, the loop you can ride for ten seconds. Do not pay a video row again hoping the next take's ambience is the Foley pack you meant.
Duration and loop are the brief
Say whether it loops. Say how long. A 10-second rain bed is not a 0.3-second click. Leaving both blank and regenerating "a bit longer" is a second generation for a parameter you could have set. Cost is per generation, shown before you confirm.
Place the file. Do not ask the SFX model to "sync to the video." Sync is an edit: the clip has a start. Nudge the start.
FAQ
Can I ask generate_music for a whoosh so I do not switch tools?
You can. You will usually get a musical phrase. Then you will still need generate_sound_effect. That is two generations where one would have done.
Should I regenerate the video when the hit is two frames off?
No. Slip the SFX clip. Regenerating picture to chase a sound you can move is a video-model bill for an edit.
Does this burn the effect into the video?
No. You get audio. Attach audio or a timeline mix puts it on the file. Generate is not lay.
Is dialogue a sound effect?
No. Speech is write and generate a voiceover. SFX is the non-speech event: hit, whoosh, footstep, ambience.