Guides

    Reading a Model's Safety Policy Before You Write the Prompt

    Refusals aren't random. The categories that reliably trigger them share one test, and knowing it lets you restructure a brief before you lose a generation.

    Versely Team7 min read

    A refusal doesn't come with a reason. The model apologizes, the generation doesn't happen, and you're left guessing which word in a forty-word brief tripped it. Most people respond by deleting a clause at random and resubmitting, which sometimes works and teaches you nothing. There's a better way in: refusals aren't random. They cluster around one predictable test, and once you can see it, you can restructure a brief before you spend a generation finding out the hard way.

    One test explains most of it

    Provider safety policies, platform disclosure rules, and AI regulation all converge on the same underlying question, even though they sit at completely different points in the pipeline: could a viewer reasonably mistake this for a real, unaltered depiction of an identifiable person, place, or event? Providers apply that test before an image or video exists, as a refusal. Platforms and regulators apply the same test after it exists, as a labeling requirement. It's the same fault line seen from both sides.

    YouTube's own creator policy names the test almost exactly. Disclosure is required when content makes a real person appear to say or do something they didn't do, alters footage of a real event or place, or generates a realistic scene that didn't actually occur — and the same policy explicitly exempts minor edits like beauty filters and color grading, plus non-realistic content like animation and green-screen effects. That's a platform stating, in writing, exactly which axis it's watching: realism plus a specific real referent, not "AI involvement" as a category on its own.

    The EU AI Act runs the identical axis into binding law rather than platform policy. Deployers of systems that generate deepfake-class image, audio, or video content carry a disclosure duty under Article 50(4) of Regulation (EU) 2024/1689 — but the obligation itself narrows the moment the content stops claiming to be real: work that is "evidently artistic, creative, satirical, fictional or analogous" gets a lighter, existence-only disclosure that doesn't have to interrupt the work to deliver it. (The deployer-facing mechanics of that obligation — what counts, how the label has to actually work, the compliance dates — are covered in full in EU AI Act Transparency: What Article 50 Means for Sponsored Content. This piece is about the step before that: what happens at the prompt, before there's any content to label.)

    Put the two together and the pattern is clear: a model refusing your brief is very often recognizing the exact shape that a platform or a regulator would require a disclosure for downstream, and declining to produce it undisclosed in the first place. It isn't squeamishness about "AI." It's the same risk assessment, run one step earlier.

    The categories that reliably trigger a refusal

    Across providers, the same handful of categories account for most refusals, and they all reduce to variations on the same test:

    • A named, identifiable real person doing or saying something fabricated. Public figures and private individuals alike — the model has no way to verify consent, and a fabricated statement attributed to a real person is the single clearest version of the harm every framework above is built to catch.
    • A specific real event recreated as if it were footage of the actual thing. News moments, real incidents, real places at a real time — restaged with invented details and presented with documentary realism.
    • Content styled to pass as authentic evidence of something invented — a fake security-camera angle, a fake news broadcast lower-third, a fake screenshot of a real interface. The realism-plus-real-referent combination again, just wearing a different costume.
    • Sexual content involving apparent minors, which every serious provider treats as an absolute, non-negotiable line rather than a judgment call.
    • Graphic violence, and use of another party's protected logo or a copyrighted character — narrower categories, but ones that show up constantly in commercial briefs without anyone intending harm.

    None of this is a mystery once you see the shared logic: the harder a brief leans on "this looks exactly like something real actually happened," the more categories it risks stacking at once.

    Restructuring a brief without losing the idea

    The fix is almost never "abandon the concept." It's swapping out the specific element that makes the brief a false claim about reality, while keeping everything that made it a good idea.

    1. Swap a named real person for a described archetype. "A firm, silver-haired executive in a navy suit addressing a boardroom" carries the same visual and narrative weight as naming a real CEO, without asserting anything about a real individual.
    2. Swap "recreate this real event" for "a scene reminiscent of." Keep the mood, the composition, the era — invent the specifics. The moment nothing in the brief claims to document an actual occurrence, the disclosure and refusal risk both drop out.
    3. Push toward stylization when the idea, not the photorealism, is the point. "Illustrated," "painterly," "stylized 3D render," and similar cues move a brief out of the photoreal-deepfake zone entirely — it's the same "obviously animated content" carve-out YouTube writes into its own disclosure policy, applied a step earlier.
    4. Use the negative prompt field on purpose, not as an afterthought. Excluding a specific real-world marker — a named brand's logo, a recognizable uniform, on-screen text styled like a real broadcast — closes off exactly the detail that tips a borderline brief into refused territory. Versely's negative prompt glossary entry covers how different models actually read that field.
    5. When you're not sure if something is a documented category or just an edge case, check Versely's safety checker glossary entry before spending a generation to find out by trial and error.

    A Versely walkthrough

    Take a common commercial brief that gets refused for stacking two categories at once:

    Before: "A photorealistic video of [a specific named celebrity] endorsing our supplement, saying it cured their chronic pain."

    That's a named real person, a fabricated first-person statement, and an unsubstantiated health claim, all in one sentence — three separate categories stacked into a single brief, and any one of them is enough to trigger a refusal on its own.

    After: "A photorealistic video of a confident fitness coach in their forties, talking directly to camera in a home gym, describing how a daily supplement fits into their recovery routine."

    Same commercial job — a spokesperson delivering a product story to camera — with none of the fabricated-real-person or fabricated-medical-claim risk attached. In Versely's agent chat, running a rejected brief through enhance_prompt and asking it to remove any named real person or unverified claim gets you most of the way to a rewrite like this one automatically; from there, Versely's model-specific prompting guides tighten the wording for whichever model you're generating on. If the concern is advertiser suitability rather than an outright refusal, that's a related but distinct axis — Versely's brand safety glossary entry covers where the two overlap and where they don't.

    Brief hygiene, not a workaround

    None of this is about getting past a filter. A brief restructured this way is usually a better spokesperson ad than the original — a real testimonial you can't substantiate was always a bigger liability than a refused generation, disclosure rules aside. Reading the shape of a safety policy before you write the prompt saves a wasted generation, and more often than not, it also catches a brief that would have caused a real problem downstream anyway, once it shipped without a label nobody remembered to add.