Guides

    The Street-Interview Format (and How to Make It With AI)

    Why the street-interview format dominates short-form and how to make it with AI: question design, editing rhythm, and a full AI vox-pop workflow without a crew.

    Versely Team8 min read

    Ask a stranger on a sidewalk what they'd do with $10,000 and you have a format that has minted more short-form accounts in the last three years than any other single structure. Street interviews — vox pops, mic-in-hand, "excuse me, quick question" — sit near the top of almost every platform's engagement charts. Media companies were built on them. Dating apps, banks, and language apps run paid versions of them.

    And here's the part that changed recently: you no longer need a sidewalk, a mic flag, or a stranger's consent form to use the structure. AI video can now produce the entire format — interviewer, respondents, ambient street, cutaway rhythm — from a script. That raises real questions about when you should (and shouldn't), which this guide covers alongside the craft of the format itself.

    Crowd of people on a city street at an outdoor event

    Why street interviews out-perform almost everything

    The format's power comes from stacked mechanics that most formats only get one of:

    • Authentic unpredictability. Viewers stay because the next answer might be anything. Scripted content can't fake that tension — but it can engineer it (more below).
    • Parasocial sampling. Each respondent is a two-second character study. Faces, accents, outfits — humans can't not watch other humans reveal themselves.
    • A built-in reply engine. Every question posed on screen is implicitly posed to the viewer. Comment sections fill with people answering the question themselves, which is the reply loop most formats have to force.
    • Infinite serialization. Same question, new people; new question, same corner. The format is a template that never exhausts, which is why accounts run one question for 200 episodes.

    For brands, the strategic value is positioning: the brand plays curious host rather than pitching. A mattress brand asking strangers "what time do you actually go to bed?" earns attention no mattress ad gets, and the product connection stays natural.

    The craft: question design decides everything

    Having watched this format closely, the difference between a vox pop that runs and one that dies is almost entirely the question. The patterns that work:

    1. Concrete over abstract. "What's the most you've ever spent on a date?" beats "what do you think about dating economics?" Numbers, specifics, confessions.
    2. Answerable in one breath. If the good answer takes 30 seconds, the format's rhythm dies. The best questions produce 3–8 second answers.
    3. A gap between what people say and what they do. "How often do you wash your sheets?" works because everyone suspects everyone else is lying. Tension = retention.
    4. Self-selecting spice. Slightly cheeky questions filter for expressive respondents. The question does your casting.
    5. Brand-adjacent, never brand-central. Ask about the problem space, not your product. The product appears in the frame, the outro, or the pinned comment.

    Editing rhythm matters nearly as much: cold-open on the single best answer (never on the question), then question card, then 4–7 answers cut tight with reaction beats, escalating to either the funniest or most heartfelt answer last. Runtime 25–45 seconds. The hook-answer you open with is doing the job hooks always do — the hooks library logic applies directly.

    The AI version: how it works

    Fully generated street interviews are now viable, and the workflow is more like showrunning than filming — you write the answers you wish you'd gotten:

    1. Script it as characters. Write 5–7 respondents with distinct voices — the over-sharer, the deadpan one, the person whose answer goes somewhere unexpected. Unpredictability is engineered in the writing.
    2. Generate consistent people. Reference-to-video models keep each respondent identical across shots — VEO 3.1 reference-to-video and Wan 2.7 hold faces, outfits, and the street setting from reference frames. Several current models generate the dialogue audio natively with the video.
    3. Sync the speech. Where a model gives you the performance but not the mouth, run the clip through AI lipsync with your voice track — accurate sync is what makes or breaks believability in this format.
    4. Assemble the rhythm. Cut exactly like the real format: best answer first, tight beats, ambient street under everything.

    Or skip the assembly: this exact format exists as a ready workflow — the NYC street interview workflow generates the full multi-scene vox pop from your question and answer beats, and workflows can run on a schedule for serialized output.

    Approach Cost per episode Time Control Authenticity
    Real street shoot Crew + day + releases 1–2 days Low (you get what you get) Maximum
    Real shoot, many episodes banked Amortized lower 1 day per 10+ episodes Low Maximum
    AI-generated vox pop Credits 1–2 hours Total Requires disclosure
    Hybrid (real A-roll, AI b-roll/cutaways) Medium Half day Medium High

    The honesty question — settle it before you publish

    Here's where I'll be direct: a generated street interview is fiction wearing documentary clothes. The real format's power partly comes from viewers believing the answers are real. So the rule I'd put in writing for any brand:

    • Disclose synthetic humans presenting as real ones. Platform policies increasingly require labeling AI-generated realistic people, and audiences punish discovered deception far harder than disclosed fiction. "Fully AI-generated — but tell me your real answer below" performs fine; getting caught performs terribly.
    • Fiction framing solves most of it. Play the format as obvious sketch comedy — absurd respondents, impossible locations, a mascot with a microphone — and the disclosure is inherent. Some of the best-performing AI vox pops leaned into it: interviews with historical figures, with pets, with the products themselves.
    • Never fabricate testimonials. An AI "stranger" praising your product is a fake review with production values. Keep generated respondents answering questions about the topic, not endorsing the brand.

    The comment section doesn't care that your respondents are synthetic if the question is real — the reply engine (people answering the question themselves) works identically, and that's most of the format's compounding value anyway.

    Making it a series

    One-off vox pops underperform their potential; this is a serial format. Practical serialization:

    • Fix the question, rotate everything else ("asking every generation the same question"), or fix the setting and rotate questions. Either gives viewers a re-subscription hook.
    • Mine your comments for next week's question. The best-performing question I've seen a brand use came verbatim from a commenter's reply.
    • Bank episodes. Generate or shoot in batches; the format's consistency is a feature, and scheduled workflow runs make weekly cadence automatic.
    • Spin winners into other formats. A question that over-performs as a vox pop usually also works as a POV scenario or a poll — one insight, three formats.

    FAQ

    What makes the street-interview format perform so well?

    Stacked mechanics: genuine unpredictability (retention), rapid-fire human character studies (watchability), and a question that viewers instinctively answer in the comments (replies). It's also infinitely serializable, which compounds an audience over time.

    Can AI really generate a whole street interview?

    Yes — reference-to-video models keep each respondent consistent across shots, several models generate speech natively, and lipsync tools handle the rest. Ready-made workflows generate the full multi-scene vox pop from a script. The constraint isn't capability; it's that synthetic humans presenting as real ones should be disclosed.

    Do I have to disclose that a street interview is AI-generated?

    If realistic AI people could be mistaken for real interviewees, yes — both platform policies and audience trust point the same way. The cleanest route is fiction framing (obviously absurd respondents, impossible settings) where the disclosure is built into the joke.

    What questions work best for a brand vox pop?

    Concrete, one-breath-answerable questions in your problem space with a say/do gap — "what's actually in your bag right now?" for an organizer brand — never questions about your product. The question should generate comment-section answers even from people who skip the video.

    How long should a street-interview video be?

    25–45 seconds for feed content: cold-open on the best answer, question card, then 4–7 tightly cut responses ending on the funniest or most human one. Longer compilations work once a question has proven itself.

    Write one great question tonight and run the NYC street interview workflow on it — or build your own cast with the AI video generator. Free credits daily.