Mix Notes Mixers Can Actually Execute
Give mixers timecode, stem, direction, amount in dB or Hz, and a reference clip. A one-page template that replaces 'make it cinematic'.
"Make it cinematic" is not a mix note. It is a mood board with no fader attached. A mixer who receives that line has to guess the stem, the time, the direction, and the amount, then spend a round proving the guess. The next round is you saying "not like that." Write the five fields instead: timecode, stem, direction, amount in dB or Hz, and a reference clip. The mixer executes. You review. That is one round, not three.
Versely will not parse this template. attach_audio_to_video in mix mode is a static pair of levels for the whole clip, not a note list. The syntax below is what you hand a human mixer, or what you follow yourself in a DAW after the generates exist. The point is that the note can be done without a meeting.
Why vague notes burn a round
A mixer can execute a change they can locate. They cannot execute a vibe.
| Note you typed | What the mixer has to invent | What usually comes back |
|---|---|---|
| Make it cinematic | Stem, time, amount, which axis (loudness, width, EQ, FX) | More reverb, more low end, a louder bed |
| The VO feels buried | How many dB, duck or EQ, for how long | Music dropped 1 dB, still masked at 300 Hz |
| Too much S | Frequency, range, split-band or wideband | A lisp, or nothing audible |
| More energy in the open | Whether that is MX, FX, or a different generate | A riser over the first line |
| Match the reference | Which 8 seconds of which file, on which stem | The whole mix pushed toward a trailer |
Each invented field is a chance to spend a revision. Ducking the music under a voiceover is a real operation with a number. "Make the voice win" is not. If you want the bed 18 to 25 dB under the voice, write that. If you do not know, say "set a relationship, then I will listen."
A second failure is notes that undo a lock. The creative brief should carry those locks. Put them on the note sheet too: "DX stays dry. Do not add verb."
The five-field syntax
One line per change. No paragraphs. If two things are wrong at the same time, that is two lines.
1. Timecode. In and out, at the frame rate of the sequence. Versely's default video frame rate is 25 fps. If the mixer is on 24, convert before you send, or write 25 fps in the header so they do not line up a pop on the wrong frame. Use HH:MM:SS:FF. A single hit can be an in-point only. A region needs in and out.
2. Stem. DX, MX, FX, AMB, or the printmaster if and only if the change is global (true peak, overall loudness). "The mix" is not a stem. If you do not know which stem, say so and ask for a solo, rather than aiming a cut at the printmaster and taking the voice down with the pad.
3. Direction. A verb a console already has: dip, lift, mute, unmute, HPF, LPF, cut, boost, narrow, widen, duck, lengthen release, move 2 frames later, remove verb. Not "fix," "improve," or "cinematic."
4. Amount. dB, Hz, milliseconds, frames, or a LUFS target. "A bit" is not an amount. Starting points are allowed if you mark them as starting points: dip ~3 dB, I will listen. A mixer can put 3 dB in and send a bounce. They cannot put "a bit" in.
5. Reference. A file, a region, or a previous approved take. ref: bed-sparse.wav 00:00:08-00:00:12 is usable. ref: that one we liked is not. If there is no reference, write none so nobody hunts Slack for a ghost.
Put the five in a fixed order so the sheet sorts. A line looks like this:
01:00:12:08-01:00:14:12 DX dip -2.5 dB none
01:00:18:00-01:00:22:00 MX HPF 120 Hz, 12 dB/oct ref: bed-sparse.wav
01:00:27:00 FX delay 2 frames later none
That is a session. "Make the middle more exciting" is not.
Amounts a mixer can execute
You do not need perfect ears to put a number on the page. You need a starting point and a willingness to listen to the bounce. These are starting points, not a broadcast spec.
| Problem you heard | Stem | Direction | Starting amount |
|---|---|---|---|
| First syllable of VO sits under the bed | MX | duck, fast attack | 9-12 dB down on speech, release slow enough that it does not pump |
| Bed is a brick in the low-mids | MX | cut | 3-6 dB, wide, somewhere in 250-500 Hz, sweep with the VO playing |
| Whole clip too hot | printmaster | lower | hit EBU R 128 true peak of -1 dBTP; do not "turn it down a bit" |
| Whoosh early | FX | delay | 1-3 frames later at 25 fps |
| Sibilance on "s" | DX | split-band de-ess | 3-6 dB in the band you swept, not a wideband dip |
| Music gone in the gaps | MX | lengthen release / raise floor | less depth, slower return; see pumping versus breathing |
Loudness is a printmaster note, last. EBU R 128 sets programme loudness at -23.0 LUFS with a true-peak ceiling of -1 dBTP. Short-form has its own supplement. Forum "targets" are not that document. Write the standard you mean, or write "match printmaster v3 integrated loudness" and attach v3.
If the generate was arranged as a record, not a bed, the note is "replace MX with a sparse instrumental," not an EQ. Adding music to a video is a static mix: music_volume versus original_volume for the whole clip. Sidechain, band cuts, and frame-accurate moves happen in the DAW. Do not write DAW notes against a tool that cannot execute them.
A one-page mix-note template
Copy this into a doc that lives next to the picture, not in a chat thread. One page. If you need a second page, you need a new mix, not more adjectives.
JOB: [name]
PICTURE: [filename, frame rate, duration]
PRINTMASTER: [filename]
STEMS: DX / MX / FX / AMB (strike any you do not have)
FRAME RATE: 25 fps unless noted
LOCKS: [e.g. DX stays dry; no added verb / MX instrumental only]
LOUDNESS: [e.g. R128 -23 LUFS / -1 dBTP, or "match v3"]
REFERENCES: [files, with timecode regions]
# TC in-out STEM DIR AMOUNT REF
1 01:00:04:00-01:00:06:10 DX dip -3 dB none
2 01:00:04:00-01:00:20:00 MX cut 4 dB @ 350 Hz, Q 1 none
3 01:00:09:12 FX delay 2 frames later none
4 01:00:14:00-01:00:18:00 MX duck 10 dB, slow release ref: v3 01:00:14-18
5 01:00:22:00-end AMB lift +2 dB none
OUT: bounce printmaster + updated stems, same duration, 2-pop intact
DUE: [date]
Fill every header field even when the table is short. A mixer who has to ask the frame rate has already spent the time you were trying to save. If a stem is missing, strike it in the header so they do not search the zip.
Send the notes with the files, the way you would hand generated clips to an editor: names that sort, one rate, a log. Mix notes are the audio half of that log.
Keep agent-executable notes separate from the DAW list. "Pull the bed under the VO for the whole clip" is an attach-audio ticket. Hz and Q are not.
Where Versely stops and the mixer starts
Add a voiceover and mix-mode music are whole-clip operations. They are the right first pass: get the relationship close, preview the picture (the editor's 480p preview is free, with a short per-user cooldown), export once. Anything that has to move in time, duck on a phrase, or cut a band happens after that export, on stems, in a tool with automation.
If you are the mixer, use the same sheet on yourself. Solo the stem, make the numbered change, bounce, listen. A 350 Hz cut and a duck on the same region interact; do them as two listens.
The review question is not "does it feel cinematic." It is "did line 2 happen." If it happened and you still hate the bed, the note was wrong. Write line 6. Do not relabel line 2.
FAQ
What if I do not know the amount yet?
Write a starting point and mark it. dip ~3 dB, starting point is executable. whatever it needs is not. Listen to the bounce, then send one number in the opposite direction if you overshot. Two numbered rounds beat one poetic round.
Should notes go on the printmaster or on stems?
On stems, unless the change is global (true peak, integrated loudness, a limiter on the way out). A printmaster dip to "make the VO sit" also dips the FX you wanted to keep. If you only have a printmaster, the first note is "we need stems," not a list of EQ moves.
Can I send this as a voice memo?
Record the listen-through, then transcribe it into the table before anyone mixes. A memo that never becomes lines is how "the bit after the logo" lands on the wrong second.
Does Versely apply these notes if I paste them into chat?
No. The agent can mix at a static level, replace a track, regenerate a line, or split a Suno song into vocal and instrumental. It does not ride a band at 01:00:18. Paste the template for humans. For the agent, write the operation it actually has: "pull the music down under this VO and export."