Coworking content is the room tone, not a SaaS explainer
Vidu Q3 I2V for diegetic room sound. Happy Horse if someone has to speak a script.
Vidu Q3 I2V for diegetic room sound. Happy Horse if someone has to speak a script. A coworking tour that explains membership tiers is a different file from one that sounds like the room. Operators ship the first because a presenter is easy to prompt. Members book the second.
The room is the product: keyboard click, HVAC, a distant espresso pull. Diegetic sound is noise the space itself makes, timed to motion. Vidu Q3 image-to-video generates that soundtrack with the picture. It is not a slide deck with a talking mascot.
The room is the soundtrack
Start with a still you already own: the common table, a phone booth, the window bay at 4pm. Image-to-video keeps that frame as opening picture. Vidu Q3 is native audio, 5 / 10 / 15 seconds, up to 1080p. Write the sound the way you write the camera:
Static wide of the long table, late-afternoon window light. Audio: keyboard clack, low conversation two tables off, espresso machine far left, no music, no speech.
Unspecified audio is how you get a music bed the room never had. The Vidu Q3 native-audio product-stories writeup is the same split: name two or three sounds, state "no music, no speech" for clean room tone, treat the mix as a starting point. Do not generate a prettier loft. A text-to-video coworking set is a room the member will not recognise at the door.
A SaaS explainer is a different file
Feature-list voiceover over stock desks is software marketing. If the brief is "hot desks, day passes" as claims, lock a real slide and typeset later. Do not invent members pointing at a pricing wall.
People in the photo are private. A generated "team collaborating" that looks like last Tuesday's occupants is a likeness problem, not a style. Shoot empty, or shoot people who signed a replica grant. Room tone does not require a generated face.
The image-to-video generator is "this still, then motion." Prompt only what should move: steam, a curtain, someone you actually photographed crossing the far background.
When someone actually has to speak
House rules, a membership FAQ, an induction line — that is a talking-presenter job. Happy Horse 1.1 image-to-video takes a first-frame still plus a script and returns 1080p with native audio and multilingual lip-sync, 3–15 seconds. Use it when the mouth is the shot.
Do not send the empty lounge to Happy Horse and hope the furniture speaks. Do not send a script to Vidu Q3 and hope the room becomes a host. Happy Horse's native-audio I2V writeup is speech-first; Vidu is scene-first. Opposite halves of native audio, not two skins of one row.
A real staff host needs consent that covers generation, not a headshot "for consistency." An invented character with no living referent still has to be disclosed. A synthetic concierge reading prices is still a claim about the space.
What to shoot before you generate
Three stills beat one prompt: long table, booth with the door ajar, street window. Same white balance. Animate each on Vidu Q3 with a different room-tone line. Cut. Caption only if someone actually spoke.
The test: would a member with a key fob say "that is our room"? If not, you made a SaaS explainer. Run Vidu on the real still. Run Happy Horse only when a script has to leave a mouth.
FAQ
Can I prompt Vidu Q3 to add a host walking through the space?
You can, and you will usually get a person the room does not employ. Keep the still empty of strangers. If a host has to speak, that is Happy Horse or a filmed staff member, cut against the room-tone plates.
Does Vidu Q3 replace recording the actual space?
No. A phone recording of the real HVAC is still the honest loop for long ambient. Vidu is for short, sound-forward social cuts from a still you already have, with diegetic noise timed to the motion it invents.
When is Happy Horse the wrong model for coworking?
Whenever the deliverable is the room, not a mouth. A 15-second talking mascot does not sell the booth. Save Happy Horse for a scripted FAQ take, then cut it after the room-tone hook.
Should I generate busy "member energy" if the photo is empty?
No. Empty is information. Fill it with sound, not with faces you do not have a grant for. Room tone is legal and cheaper than a replica you cannot defend.