Turning sound on switches the whole clip's per-second rate
On Veo 3.1 the audio toggle replaces the silent rate for every second; attaching a separate VO is additive instead of multiplicative.
On Veo 3.1 the audio toggle replaces the silent rate for every second; attaching a separate VO is additive instead of multiplicative.
Native audio is not a surcharge sitting on top of a mute bill. It is a two-value lookup. VEO 3.1 bills 20 credits a second silent and 40 with sound. The formula on AI video with sound is credits = rate_per_second(audio on | audio off) × seconds. Flip the toggle and every second of the clip moves to the other rate. There is no line item for “the talking part.”
That is the whole pricing object. Mix it up with “I’ll just add a voiceover later” and you will double-pay or under-quote.
The toggle replaces the rate
On VEO 3.1 an 8-second shot is 160 credits mute and 320 with native audio. The premium is 160 credits for that length — 20 credits a second, charged on all eight seconds, not on the three seconds someone speaks. The same toggle at a minute of that rate would be 1,200 extra credits. Multiplication grows with length. That is why native audio wins on very short clips and loses on long ones.
VEO 3.1 sells fixed durations: 4s, 6s, 8s. You do not shave a second to save a little. You pick a block, then you pick whether that block is silent or sounded. The toggle — not the length — is the lever you actually operate.
Nine models in the catalogue bill this way (audio_based / audio_based_per_second). The multiple is not a constant. VEO 3.1 is 2×. Other rows are gentler or steeper. Best model with native audio is the capability ranking — models whose category or feature tags actually name audio, not the raw audio boolean, which is set on silent video rows and does not mean a priced toggle. Capability and the meter are different pages.
What native audio buys that the additive route cannot is synchronisation: sound generated with the picture is matched to what is happening on screen. That is a creative argument. Pay for it when you need lips, foley, and room with the frame. Do not pay for it because the UI had a switch.
The whole clip, not the talking part
The rate is replaced for every second. A clip that is silent for five seconds and voiced for three still bills eight seconds at the audio-on rate if the toggle is on. There is no prorating. The model is generating picture-with-sound as one object.
Do not:
- Turn audio on “just for the line” and expect mute pricing on the pads.
- Compare an 8-second native-audio quote to a 4-second silent quote and call the difference “the cost of sound.” You changed two variables.
- Assume 4K on VEO 3.1 is a third switch on the same meter. Resolution is a different lookup. Sound is the audio dimension.
If the job is a talking close-up, native audio is the honest row. If the job is a product orbit that will wear a separately written VO, generate mute and attach. The cost hub exists so you can see which meter a real generate will hit before you dispatch it.
Additive voiceover is a different meter
The other way to get sound onto a clip is to generate the picture mute, generate speech as its own item, then attach. Speech bills per 1,000 characters of script, rounded up. The step that lays the file onto the video is a flat attach fee. The video’s own price never moves. Length does not grow the attach. One voiceover can cover several clips.
That route is additive: mute video + TTS + attach. It behaves nothing like the toggle. It is usually cheaper on anything but the shortest native-audio blocks, and the audio is reusable. It is also not lip-synced native sound. You are paying for a read, not for the model to speak with the face.
Voiceover for a script is that character meter. Add a voiceover to a video is the attach job. Music is the same additive shape: generate a bed, attach, flat fee, one track can cover several clips. Dubbing an existing video’s dialogue into another language is a different job again — a flat per-job charge, not this toggle.
Do not add native audio and a full VO on the same take unless you meant to. You will have paid the 2× rate for sound the attach is about to cover.
Draft silent, pay for sound once
Silent-first drafting is the obvious saving. Every mute 8-second VEO 3.1 attempt is 160 credits instead of 320. Iterate the picture mute. Turn the toggle on once the frame is the one you will ship. The premium is then paid once instead of on every miss.
That only works when native audio is not the thing you are testing. If the job is “does the line land with the mouth,” mute drafts teach you nothing. Pay the audio-on rate for those takes. If the job is “is this the kitchen,” mute is enough.
The toggle is never charged on only part of the clip. Plan the duration block first (4 / 6 / 8), then the sound setting, then the number of retries you can afford at that combined rate. Quote it on the cost page before you confirm.
FAQ
Is native audio a surcharge on top of the video price?
No. It replaces the per-second rate. On VEO 3.1 silent seconds bill at 20 credits and audio seconds at 40. There is no separate line item, and the audio-on rate applies to every second of the clip.
Is it cheaper to add a voiceover instead?
Usually, on anything but the shortest clips. A voiceover is billed per 1,000 characters and the attach step is a flat fee that does not grow with duration. You lose native sync. You gain a reusable read.
Can I turn audio off to iterate and on for the final render?
Yes, and it is the obvious saving — every silent 8-second draft is 160 credits instead of 320 on VEO 3.1. Use it when you are iterating picture. Do not use it when the test is the sound.
Do all video models charge for audio this way?
No. Nine models expose audio as a priced toggle; the rest are silent by design and rely on the additive route. The multiple is not the same on every toggle row. Read the rate pair on the model you are actually running.