Keep Finishing Audio at 48 kHz
Resample every 44.1 kHz generate on ingest, use a real anti-alias filter, and keep the project at 48 kHz / 24-bit so export does not click.
Step-by-step guides for making video, images, voiceovers and music with AI — written for people shipping content, not reading theory.
Page 41 of 62
Resample every 44.1 kHz generate on ingest, use a real anti-alias filter, and keep the project at 48 kHz / 24-bit so export does not click.
Generated music parked in 5.1 L/R starves dialogue on stereo TVs. Fold down with center protected and LFE discarded for social and web masters.
Generated clips arrive display-referred. Read Rec.709 waveform, RGB parade and vectorscope, then grade exposure, balance, sat, look.
Models return exact-length clips. Pad 8-12 frames on each end so dissolves, J-cuts, and limiter look-ahead do not hit hard cuts.
An HI mix is not louder voiceover. Dialogue-forward EQ, reduced music, and a stem recipe for shipping HI and VI tracks instead of relying on captions alone.
Grade Rec.2020 PQ or HLG once, roll off highlights, and QC two stills of the same frame so the SDR YouTube downmap does not crush the HDR master.
Brand sites still ship one 1080p MP4. A six-rung HLS and DASH ladder, aligned keyframes, CMAF, and loudness so the owned player matches the social encode.
Hume Octave 2 leads on emotional fidelity in text-to-speech. When that's worth the premium over a flat voice, and when it's genuine overkill.
Netflix still wants an IMF package, not an MP4. Walk the CPL, PKL, markers, and supplemental versioning, and what a small studio can build without an IMF house.
Space, arrows, a captions toggle, and a visible focus ring. Why marketing embeds fail, and a 12-key staging test you can run in ten minutes.
Models often write full-range RGB into Rec.709 tags. Legalize after the grade, never before, and confirm 16-235 on a waveform before delivery.
Hold lips to ±1 frame on film and ±2 on broadcast. Slip generated SFX inside that window, and regenerate only when a slip would break the rest of the scene.
Use published integrated, true-peak, and LRA figures per destination, and keep two masters when YouTube, Netflix, BBC, and podcasts disagree.
Scrim, stroke, and plate choice so a name super holds on generated bokeh and neon. A three-treatment test you run on a 480p preview before export.
Spoken Chinese is Mandarin or Cantonese; written Chinese is simplified or traditional. Pick all four, and never ship a Hong Kong cut as Taiwan.
Laptop speakers lie about bass and sibilance. Spec a small-room chain: monitors or headphones, a measurement mic, 83 dB pink, and a phone-speaker pass.
Mid-side generated music can cancel in mono on TVs and some phones. Run a ten-second fold, read a correlation meter, and collapse width without killing the bed.
English supers overflow in German and Finnish and starve in CJK. Set expansion budgets, two-line fallbacks, and a regenerate-the-plate rule before type shrinks.
Rapid generated cuts and lightning VFX trip Harding and ITU-R BT.1702. A three-flashes-per-second test, a flashmeter pass, and recuts that do not regenerate.
Center-bottom captions sit on generated mouths and burned-in prices. WebVTT and TTML regions, when to raise or split, and a vertical safe-zone review.
Generation QA grades a clip against its prompt. Finishing QC signs off freeze frames, black, sync, flash, loudness, captions, slate, textless, and a checksum.
Generated faces sit off the skin-tone line. Read the vectorscope, then use hue and sat moves to unify a sequence without running a beauty pass.
Model defaults love lightning, glitch, and strobe. Photosensitive thresholds, a pre-publish flash test, and how to dim or recut for web and social.
Social files skip leader. Agency and broadcast still want eight seconds of bars and tone, a 2-pop, and a slate listing job, version, LUFS, and duration.