Guides

    Keeping reverb out of generated dialogue

    Reverb is the most common uninvited guest in native audio. The room words that cause it, the exclusion phrasing that works, and what to do with a good take.

    Versely Team9 min read

    Reverb is the one thing in a generated soundtrack you cannot fix later. Every other audio problem has an escape hatch: a bed that is too loud gets pulled down, a wrong word gets re-read, a missing effect gets layered in. Reverb is convolved into the voice signal itself. Once the model has decided your presenter is standing in a stairwell, that stairwell is part of the waveform.

    It is also the most frequent uninvited guest in native audio output, which is an awkward combination. Here is why it happens, the vocabulary that triggers it, and the phrasing that keeps it out.

    Why models add reverb you did not ask for

    Nothing in the model wants reverb specifically. What it wants is plausibility, and reverb is what plausibility sounds like.

    Generated speech is conditioned on the scene you described. Describe an interior and the model reconstructs what a voice in that interior would sound like, which includes early reflections off the walls you implied. That is correct behaviour — a voice recorded in a warehouse with no reverb sounds fake, and the model has learned that. The problem is that it does not know you intended to cut this clip against three others recorded in different implied rooms, or to lay a clean voiceover over the top, or to run the line through captions where intelligibility beats realism.

    So the default is wet, and wet is the wrong default for almost everything short-form. Phone speakers already smear transients; a reverb tail on top of that turns consonants to porridge. A line that reads fine in headphones can be unintelligible in a feed.

    There is a second, sneakier source: implied distance. If the shot description puts the speaker at the far end of a room, or the camera on a wide, the model matches the acoustic to the framing. Wide shot, distant voice, more room in the signal. This is one of the few places where cinematography language leaks straight into the audio, and it catches people out because nothing in the prompt mentioned sound at all.

    The words that summon a room

    Certain descriptions reliably produce reverberant dialogue. None of them are audio words.

    Prompt language What arrives on the voice
    "warehouse", "hangar", "cathedral", "gymnasium" Long tail, obvious slap
    "empty apartment", "bare room", "concrete stairwell" Hard early reflections, boxy
    "hallway", "corridor", "underground car park" Flutter echo, comb filtering
    "the camera is far away", "wide shot of the room" Distance-matched, thin and washy
    "her voice echoes", "she calls out" Deliberate reverb, requested by accident
    "tiled bathroom", "swimming pool", "marble lobby" Bright, splashy, high-frequency ring

    The last row is worth calling out because tile and marble are visual choices. Someone picks a marble lobby because it looks expensive, and gets the lobby's acoustic along with it.

    The counter-list is short: carpeted interiors, cluttered rooms, soft furnishings, cars, small offices, curtained spaces, and anything outdoors that is not a canyon. Outdoors is the reliable one — open air has nothing to reflect off, so any wash you hear there is wind rather than reverb.

    Exclusion phrasing that actually works

    The instinct is to write "no reverb". It half-works, and understanding why tells you what to write instead.

    Instructive negation — "no reverb", "don't add echo", "avoid room sound" — puts the negated noun in the prompt. Models are inconsistent about honouring the negation and consistent about noticing the noun, which is why the pattern sometimes produces exactly what you excluded. This is the same failure that makes negative prompts unreliable on modern architectures, and the same fix applies: describe the state you want rather than naming the state you don't.

    Three phrasings, in ascending order of reliability:

    1. No reverb. — a coin flip.
    2. Ambient noise: a small, acoustically dead room. — describes a state, mostly works.
    3. Close-mic'd dialogue, dry and intimate, recorded in a small carpeted room with soft furnishings. — describes a microphone position, an acoustic, and a set of materials. This is the one that holds.

    The third version works because it gives three independent signals that all point the same way. "Close-mic'd" implies proximity, which suppresses distance cues. "Dry and intimate" is a recording description rather than an instruction. "Carpeted, soft furnishings" gives materials that physically absorb. Any one of them alone is weak; together they leave the model very little room to add a tail.

    Two supporting moves:

    Shorten the shot size. A close-up or medium close-up carries an implied proximity that a wide does not. If the line matters more than the framing, take the closer shot — you are buying acoustic dryness with the same token that sets the frame.

    Put the exclusion in its own sentence. Audio instructions stacked into one clause get parsed as a single blurred request. One idea per sentence is the rule that governs the whole audio side of a prompt, and it applies to exclusions as much as to the effects you name — scripting dialogue and sound cues covers the wider syntax.

    A prompt that behaves:

    Medium close-up, static camera. A man in a grey jacket looks off-camera and says, "We're not waiting for them." Close-mic'd dialogue, dry and intimate. Ambient noise: a small carpeted room, soft furnishings, no music.

    When the take is good and the room is wrong

    Sometimes the performance lands and the acoustic doesn't. There are three honest options, and one of them is not "de-reverb it in Versely" — vocal isolation strips background music and instrumentation from a mixed signal, which is a different job from removing a room from a voice.

    Option 1: re-roll. Cheapest when the clip is short and the performance was not the point. Adjust the room description, take the closer shot, and generate again. Most reverb complaints resolve here.

    Option 2: replace the voice entirely. If the picture is what you wanted and the audio was incidental, generate the line as clean speech with text-to-speech and swap it in. Replacing a video's audio strips the original track and substitutes yours, the video stream is copied without a visual re-encode, and the fee is flat regardless of clip length. You lose native lip sync accuracy on a talking face, so this works best on off-screen lines, B-roll narration, or shots where the mouth is not the subject. The full comparison of native audio against the TTS path covers when that trade is worth making.

    Option 3: dedicated de-reverberation, outside the pipeline. Restoration suites in the iZotope RX class have de-reverb processing built for exactly this. It is a real fix, not a miracle: light room sound comes off cleanly, a genuine stairwell tail does not, and aggressive settings leave the voice sounding hollow. If the take is irreplaceable — a performance you cannot reproduce, a multi-speaker exchange that took eight attempts — this is the path. For anything you can simply generate again, it is more work than re-rolling.

    All three cost more than getting the room description right the first time, which is the argument for spending an extra clause on acoustics in every prompt with a spoken line in it.

    Checking before you commit

    Play the clip on a phone speaker at arm's length, in a room with some noise in it. If consonants survive, you are fine. If the ends of words disappear, there is a tail on the voice regardless of what your headphones said. Do this on a 480p preview rather than a paid export — audio comes through the preview pass unchanged, so it answers the question completely, and previews carry no credit cost beyond a short per-user cooldown. The charged render is the final export, once, no matter how many clips are on the timeline.

    The other check is comparative. Cut two candidate takes back to back and listen for the moment the room changes. Individual reverb is tolerable; reverb that changes at every cut is what makes a stitched sequence sound assembled rather than shot. Fixing that at the assembly stage is a separate job, and it starts here, with each clip coming out drier than feels natural in isolation.

    FAQ

    Does "no reverb" in the prompt ever work?

    Sometimes, which is the problem. Instructive negation leaves the word in the prompt and depends on the model honouring the negation, so results vary between runs on the same prompt. Describing the acoustic you want instead — close-mic'd, dry, small carpeted room — removes the coin flip. Use the exclusion as a backstop after the positive description, not as the whole request.

    Why does my dialogue get more reverberant on wide shots?

    Because the model matches the acoustic to the implied distance between camera and speaker. A wide shot suggests the microphone is far away, and a far microphone picks up more room than direct sound. If the line has to be clean, take the closer framing, or state the mic position explicitly so the acoustic follows the mic rather than the lens.

    Can Versely remove reverb from a clip I already generated?

    Not directly. Vocal isolation removes background music and instrumentation from a mixed recording, which is a different problem — isolating vocals from a track will not pull a room out of a voice. The realistic routes are re-rolling with a drier room description, replacing the audio with generated speech, or running a dedicated de-reverb process in a restoration suite before bringing the cleaned track back in.

    Is a little reverb ever the right call?

    Yes, when the scene is meant to feel like a place rather than a read — a character speaking in a church, a shout across a car park, anything where the space is part of the story. The rule is about unintended room sound on lines that carry information. Narration, product copy and calls to action should be dry, and dry is not what the model gives you unless the prompt asks for it.