Guides

    AI music with a brittle, glassy high end

    Generated tracks carry a glassy 4–7 kHz sheen that bites on phones. The band to cut, how much, and why a mix pass beats another prompt.

    Versely Team8 min read

    Play a generated bed on studio monitors and it can sound expensive: bright, wide, plenty of air. Play the same file on a phone speaker, a laptop, or a cheap earbud, and the top turns to glass. Cymbals become a hiss. Vocal esses stick. Pads get a metallic edge that was not in the room, because there was no room. That upper-mid sheen is a mix problem with a named band. Prompting "warm analog, no harsh highs" does not reliably remove it, because the model is not mixing. It is emitting a finished stereo file that already has the brightness printed in.

    The fix is a cut in the 4–7 kHz range on the music, a few dB, checked on a small speaker. Do it on a WAV. Do it before you duck the bed under a voice. Do not try to generate the sheen away.

    The 4–7 kHz fingerprint

    Human speech intelligibility lives in a neighbouring region, which is why this band is so unforgiving in a video mix. Presence, sibilance, pick attack, and hat splash all pile up between about 4 and 7 kHz. Generated arrangements tend to put all of those at once, at similar levels, with no microphone distance and no air absorption to roll the top off. The result is not "too much treble" in the 10–12 kHz sense. It is a hard, glassy presence peak. Boosting a high shelf to add "air" makes it worse.

    Two related problems get misdiagnosed as the same thing:

    • Mud is 250–500 Hz: voice, bass harmonics, pads, and guitar body stacked. The track feels thick and the words get covered. Carve that on the bed if the voice is fighting for space. It is not the glassy bite.
    • Sheen is 4–7 kHz: the track feels sharp, especially after a few seconds, especially on small speakers. That is this post.

    You can have both. Generated full mixes often do. Fix them in that order: mud first so the voice has a body, sheen second so the voice does not get cut to ribbons by the hats. Always work from a WAV of the generated track. An MP3 has already thrown away the top you are trying to shape.

    Why phones make it worse

    Phone speakers are tiny, peaky, and weakly extended at both ends. They cannot reproduce sub or true air, so they dump what energy they have into the presence region. A 4–7 kHz build that was merely "crisp" on a full-range monitor becomes the entire personality of the playback. Earbuds do a milder version of the same thing: they couple close to the ear and make sibilance and hat splash feel closer than they were in the mix.

    That is why the check is not your nearfields. After you make the cut:

    1. Bounce a 15-second loop that includes a vocal or a hat pattern.
    2. Play it on the phone speaker at a comfortable talking volume, not loud.
    3. Play it again with a voiceover sitting on top, because the sheen is worst when it shares the band with spoken sibilance.

    If the loop is tiring after one pass, the cut is still too shallow. If the track suddenly sounds like it is coming through a blanket, you went through the body of the music, not the sting. Narrow the band or reduce the cut.

    The cut: band, width, how much

    A starting recipe that survives contact with generated beds:

    Move Setting Why
    Find the sting Sweep a wide bell through 4–7 kHz at +4 dB until it hurts, then flip the gain to a cut You are locating the sheen, not guessing 5.0 versus 6.3
    Cut 2–4 dB, Q around 1.0 (wide) A notch punches a hole; a wide bell tames a region
    Optional air A small high shelf above ~10 kHz, 1 dB or so, only after the 4–7 kHz cut Restores a sense of space without putting the sting back
    Do not Boost 4–7 kHz, or add a "presence" preset on the master That is the band you are trying to leave

    2–4 dB is a starting amount, not a measured law. Dense hat-heavy beds often want the top of that range. A sparse acoustic bed may only need 2 dB, or may only need the cut on a hat/cymbal stem if you have one. Make the smallest cut that passes the phone-speaker test.

    A few constraints that actually matter:

    • Cut the music, not the voice. If you EQ the voice down at 5 kHz to dodge the bed, you dull the words and keep the hats. The bed is the offender.
    • Prefer a bell in 4–7 kHz over a low-pass. A low-pass at 8 kHz kills air and still leaves the sting if the peak is at 5 kHz.
    • If you have stems, you can put the cut on hats, cymbals, or a bright vocal stack instead of the whole mix. Treat Suno stems as separated from a mixture, not as isolated mics. Ghost vocals on the drum stem and reverb tails on the vocal stem are normal. A wide 4–7 kHz cut on the stereo bed is often cleaner than pretending the stems are a recorded session. Stem separation is still worth doing when a sung hook is competing with your voiceover; mute the vocal stem rather than EQing a ghost out of the drums.

    Work on the WAV from text-to-music generation, then attach the treated file. Do not print the EQ into a low-bitrate bounce and then mix that under a voice.

    Why another prompt will not remove it

    The sheen is a spectral habit of the finished render, not a lyric or a genre token the model is waiting to hear. "Warm", "analog", "dark", "no harsh cymbals", "vintage" will sometimes give you a different arrangement. They will not consistently roll off a 5 kHz build that the decoder likes. You will burn credits chasing a spectral shape the mix window can change in twenty seconds.

    Prompt for musical content: genre, instrumentation, energy, whether you want an instrumental. Then mix. That split is the same one that shows up across generated media: most of the defects that survive first listen are cheaper to fix in post than to regenerate. AI music for brand videos is the generation side. This post is the EQ you run after you already have a track you like.

    A prompt that is still worth writing, because it changes arrangement density rather than EQ:

    Instrumental, mid-tempo indie electronic, soft analog synths, brushed kit, no bright crash, no sung vocal, spare arrangement, room for a spoken voiceover

    That can give you fewer hats to fight. It will not replace the 2–4 dB cut once a bright render lands.

    Put the treated bed under the voice

    Once the sting is gone on a phone speaker, the rest of the mix is ordinary:

    1. Keep the treated file as WAV.
    2. If the bed is shorter than the picture, extend the music track instead of looping a bright eight-bar phrase into a seam.
    3. Add the music to the video in mix mode, with the bed well under the voice. If you also need a moving duck, print that duck on the already-EQed WAV. Ducking a brittle bed just pumps the sheen in and out of 4–7 kHz, which is more tiring than a static bright track.
    4. If a sung vocal is still occupying the same presence band as the voiceover, split the song into stems and use the instrumental.

    Generate from the AI music generator. Mix as if the file is a stereo bounce from someone else’s session, because that is what it is. The band is 4–7 kHz. The amount is a wide 2–4 dB to start. The check is a phone speaker with the voice on top. The prompt is not the equalizer.

    FAQ

    Where exactly should I cut?

    Sweep a wide bell through 4–7 kHz until the harshness jumps out, then cut 2–4 dB there. Generated beds do not all peak at the same frequency inside that range. A fixed 5 kHz notch will miss a 6.5 kHz hat build and still leave the bite.

    Will a "warm" prompt remove the sheen?

    Sometimes you get a darker arrangement. You do not get a reliable 4–7 kHz cut. Prompt for instrumentation and density, then EQ the WAV. Regenerating to chase a spectral shape is how a one-minute mix pass turns into an afternoon of credits.

    Should I cut the voiceover at 5 kHz instead?

    No. The voice needs that band to be intelligible, especially on phones. Cut the music. If spoken sibilance is a separate problem, de-ess the voice on its own. Do not use the voice as a dump for the bed’s hats.

    Can I just low-pass the whole track at 8 kHz?

    You will lose air and often keep the sting, because the sting is usually below that filter. A wide bell in 4–7 kHz, 2–4 dB, is the move. Add a tiny shelf above 10 kHz only if the bed feels dull after the cut, and only on a WAV.