The perception gap between you and your audience
Ad professionals and younger consumers are 37 points apart on AI ads, and the gap widened. A per-brand tolerance test you can run in a fortnight.
The number most people in this industry watch is the one on the invoice: will clients ask for a discount now that production is faster. The number that should actually worry them is the one measured on the other side of the screen. eMarketer reports a 37-point perception gap between advertising executives and Gen Z and millennial consumers on AI-made ads — and it widened from 32 points in 2024.
Widening is the part that matters. The comfortable assumption was that exposure would normalise generated advertising the way it normalised stock photography and CGI. Two years of data point the other way: as the volume went up, the distance between how the industry sees this work and how younger audiences see it got larger, not smaller.
The industry is not unaware, it is differently focused
The same body of research shows 60% of US ad professionals naming accuracy and transparency as a top barrier to AI adoption. So the concern exists inside the industry; it is simply pointed at a different object. Professionals worry about whether the output is correct and whether the disclosure is compliant. Audiences are reacting to something upstream of both — a judgement about effort, care and whether the brand meant it.
Coca-Cola's holiday campaign is the clearest public instance of the two frames colliding. On the production side it was a success on its own terms: a timeline that had historically run about a year compressed to roughly a month, with the CMO confirming it came out faster and cheaper. On the audience side it drew significant public backlash. Both of those are true simultaneously, and a shop that only tracks the first will keep being surprised by the second.
Population averages do not tell you about your brand
A 37-point gap measured across "ad executives" and "younger consumers" is a distribution, and your brand sits somewhere on it that you cannot infer from the headline. The plausible movers are not mysterious, but they are brand-specific enough that they have to be measured rather than assumed:
- What the product promises. A category whose value proposition is human care, craft or provenance carries more exposure than one whose proposition is convenience or price.
- Where the AI shows up. Generated b-roll behind a real person reads differently from a fully synthetic spokesperson. What actually converts with AI avatars in ads is the narrower version of that question.
- Format expectations. Audiences hold a polished brand film and a UGC-style clip to different standards of authenticity, which is why UGC-style ads versus polished ads is not just an aesthetic choice here.
- Age. The gap in the reported data is specifically against Gen Z and millennials. If your buyer and your product's user are different generations, you may have two answers.
None of that is measurable from a trade headline. It is measurable from your own creative, which is the entire point of what follows.
A tolerance test you can run in a fortnight
The goal is not a statistically clean study. It is a directional read on where your specific audience sits, refreshed often enough to stay current.
- Pick one claim and hold it fixed. Same script, same length, same hook, same call to action across every cell. If the copy changes, you are testing copy.
- Build three executions of it. Fully generated. Hybrid — generated environments and b-roll, real hero performance or real voice. And a filmed or archive-footage version as the reference point, if you have one available.
- Disclose all three. Partly because labelling obligations are now dated and enforceable, and partly because an undisclosed test measures the wrong variable: you would be measuring whether people notice, when the question is how they respond when they know. Writing an AI disclosure line nobody scrolls past covers the wording, and profile-level versus post-level disclosure covers where it goes.
- Deliver paid, not organic. Organic distribution introduces an algorithmic variable you cannot hold constant. Paid delivery to a defined segment is the only way the three cells see comparable audiences. The mechanics are in A/B testing video creative properly.
- Measure three things. Hook rate, completion, and a sentiment read — a post-view brand question if you have panel access, comment and reply sentiment if you do not. Do not use conversions alone; a spot can convert on price while damaging the brand line, and that is precisely the trade this test exists to expose.
- Segment the read by age band. The published gap is generational. Collapsing your result to a single number reproduces the exact error the headline number invites.
| Execution | What it tells you |
|---|---|
| Fully generated | Your audience's floor — the most exposed version |
| Hybrid | Whether a visible human layer recovers the difference |
| Filmed reference | The ceiling you are trading against, in your own numbers |
Run the same three cells again at the end of the quarter. A gap that widened industry-wide over two years is not a constant you measure once.
What a bad result actually licenses you to change
The reflexive conclusion — stop using AI — is not what the credible cases support. The pattern across the campaigns with real corroboration is narrower and more useful: generative tools carry exploration and volume, human craft carries the finish. Nike's work with Serena Williams generated 130,000 virtual tennis matches as an exploration step, with human craft applied to what shipped. Klarna reported roughly $10M a year saved on external agencies and content. Neither of those is a story about removing people from the output.
So when a cell underperforms, the change is usually about where the human layer sits rather than whether the tooling is used at all:
- Put a named, real voice on it. Voice work is the cheapest visible human layer in most ads.
- Keep the hero performance real and generate everything around it.
- Make the writing carry more. A script that could only have been written by someone who knows the category is the least imitable thing in the ad.
- Fix the finish. Grade, sound design and cut rhythm are where "generated" usually announces itself, and they are all edit-side decisions.
There is a capacity dimension to this too, and it cuts against the assumption that volume solves the problem. Superside's Breakpoint research found 80% of creative teams at or beyond capacity and 70% of creative leaders burnt out — despite AI adoption. Generation stopped being the bottleneck; filtering, governance and taste became it. Producing four times as many variants does not help if the constraint is the judgement applied to them, and audience tolerance is a judgement problem.
The conversation that is not the risk
Worth saying plainly, because it consumes attention the audience question deserves: the client-side pushback is smaller than the industry fears. Reported survey figures put 73% of agencies as never having been asked to cut prices despite adopting AI, and of the 27% who were asked, only 13% actually lowered rates. Clutch's 2024 data has 61% of agency clients raising AI during renewals — but raising a topic and demanding a discount are different events, and the follow-through numbers say so.
Which reframes the whole thing. The commercial risk in generated advertising is not that your client will pay you less. It is that your client's customers will respond to the work differently than your client expects, and nobody in the room has measured it. What to tell clients about using AI is the conversation; the tolerance test above is what turns it from opinion into evidence.
FAQ
Should we disclose even where we are not legally required to?
For a test, yes, unambiguously — an undisclosed cell answers a different question. For live campaigns it is a brand decision, but the obligations have hard dates now and the practical posture most teams land on is disclose consistently rather than case by case. An AI disclosure policy for companies is the framework for setting it once.
Isn't this just going to normalise over time?
That was the assumption, and the reported movement from a 32-point gap in 2024 to 37 points is evidence against it, at least so far. It may still normalise. Planning on it is a forecast, not a finding, and the test above costs less than being wrong about it for a year.
How large a sample do we need for a valid read?
You will not get statistical power from a two-week creative test, and you should not pretend otherwise. Treat it as directional: a consistent ordering across two segments and two time periods is worth acting on, a single narrow win is not. What makes it useful is repetition on a schedule, not sample size in any one round.
Does the gap apply to B2B?
The published figure is consumer-facing and generational, so applying it directly to a B2B buyer is unsupported. The mechanism — an audience judging effort and care — plausibly transfers, but that is an inference. If it matters to your budget, run the three cells against your own segment rather than borrowing a consumer number.