Strategy

    Calculating ROI on AI-Generated Content

    Calculating ROI on AI-generated content: the formula, what to hold constant, attribution traps, leading indicators, and how to prove capacity gains.

    Versely Team8 min read

    A performance team told me their AI content ROI was 900%. The math was: tool spend for the quarter, divided into the revenue attributed to every paid social campaign that ran in the same period. Every campaign, regardless of whether the creative was generated or shot, regardless of what the media budget was doing, regardless of the seasonal lift they get every year in that window.

    That number is worse than useless, because the first finance person to look at it will find the flaw in ninety seconds and then discount everything else you say. The irony is that the real ROI on their program was excellent — it just wasn't 900%, and it wasn't a revenue story at all.

    Getting this right requires deciding what question you're answering before you pick a formula. There are three different ROI questions people conflate, and they need three different measurements.

    Marketing team analyzing content performance data on a monitor

    The three ROI questions, and which one you're actually asking

    Question 1: Did we produce the same output for less? This is a substitution question. Hold output constant, compare cost. It's the easiest to measure and the easiest to defend.

    Question 2: Did we produce more output for the same money? A capacity question. Hold spend constant, compare volume and its downstream effect. Harder, but usually where the real value sits.

    Question 3: Did the content perform better? A creative-effectiveness question. This one requires holding almost everything else constant and is the hardest to answer honestly — but it's also the one that produces the biggest numbers, which is why people reach for it first and get burned.

    Most teams should answer 1 and 2, and treat 3 as a separate experiment. Blending them produces the 900% figure.

    The substitution formula

    For question 1, the arithmetic is simple and the discipline is in the inputs:

    Cost saved = (Baseline cost per published asset − New cost per published asset) × assets published

    Baseline cost per asset must be fully loaded: freelance and agency invoices, stock subscriptions, shoot costs, software, and internal hours at loaded rates. New cost per asset is tool consumption plus internal hours. If you only count invoices on the old side and only count credits on the new side, you'll produce a savings number that collapses under the first question.

    Two adjustments that keep this honest:

    • Count published, not produced. Assets that never shipped are cost on both sides.
    • Include the review step. Review time usually goes up per asset in the first quarter, because reviewers are checking things they used to trust. It comes back down. Reporting the first quarter's number as steady-state understates your program.

    The capacity formula

    Question 2 is where most of the value lives and where most teams stop measuring, because "we made more stuff" doesn't feel like an ROI figure.

    Make it one by valuing the incremental output at what it would have cost to produce the old way:

    Capacity value = incremental assets published × baseline cost per asset

    If you shipped 40 videos a month where you used to ship 12, the 28 extra assets are worth 28 × your baseline cost — not because you'd have spent that, but because that's the market rate for what you now have. Label it clearly as avoided cost of equivalent output, not as savings. Finance readers accept that framing; they reject it when it's presented as cash.

    Then, and this is the part that gets skipped, show what the incremental output did. Extra assets have to earn their existence:

    Incremental output Downstream metric to report
    More paid creative variants Winning-variant discovery rate; cost per acquisition trend
    More organic posts per week Reach and follower growth vs. prior cadence
    Localized versions Engagement in the target market vs. the English original
    Assets for previously unserved channels Whether the channel produced anything at all

    If the extra 28 assets moved none of these, the honest conclusion is that you have surplus capacity, not surplus value — and the right response is to redeploy it, not to inflate the number.

    Measuring creative effectiveness without fooling yourself

    Question 3 needs a controlled comparison. In practice:

    • Same campaign, same audience, same budget, same time window. Split the media budget between generated and conventionally produced creative. Anything else and you're measuring seasonality or targeting.
    • Enough volume to matter. Creative testing needs meaningful sample sizes per variant. Small tests produce confident-sounding noise.
    • A stable success metric. Pick one — cost per acquisition, thumbstop rate, completion rate — before you run it. Choosing afterward from a list of six is how every test "succeeds."

    The result you should expect, honestly, is parity on per-asset performance and a large advantage on portfolio performance. Individual generated assets rarely outperform a good shoot. But testing 20 variants instead of 3 finds a better winner, and that's a real, measurable, defensible edge — it just belongs in the capacity bucket rather than the effectiveness bucket.

    Leading indicators to track weekly

    Quarterly ROI is a lagging measure. These four tell you in week three whether the quarterly number will be good:

    1. Usable assets per human hour. The core capacity metric. Should climb steeply between weeks two and five as prompts and workflows stabilize.
    2. Keep rate. Usable ÷ generated. Flat keep rate after six weeks means a prompting or process problem, and it's the single biggest driver of cost per asset.
    3. Cycle time, brief to live. What the rest of the business feels.
    4. Tier mix. Share of generations on premium vs. fast models. Drift toward premium doubles cost at flat output, quietly.

    Deliberately excluded: total generations, credits consumed, prompts written. Activity metrics look like progress and prove nothing.

    Four attribution traps

    • The whole-channel trap. Crediting all paid social revenue to the tool that made some of the creative. This is the 900% error. Attribute only the delta you can isolate.
    • The seasonality trap. Comparing a Q4 program to a Q2 baseline. Compare like periods, or year over year.
    • The survivorship trap. Measuring only the generated assets that shipped, against all conventional assets including the duds. Keep rate belongs on both sides.
    • The free-labor trap. Claiming savings on internal time that nobody reallocated. If the hours didn't go anywhere, they're not saved — they're idle, and someone will notice.

    Reporting it so it survives the room

    One page. Three sections in this order:

    Substitution. "We published the same 12 monthly explainers at $X per asset, down from $Y." Small, verifiable, bankable.

    Capacity. "We additionally published 28 assets per month that previously would not have existed, at an equivalent production value of $Z. Those assets produced [specific downstream result]."

    Effectiveness. "In a controlled split on the June campaign, generated creative performed within N% of conventional on cost per acquisition, and the wider variant pool improved the winning creative's CPA by M%."

    Then the caveats, unprompted. A section that lists what you couldn't measure buys more credibility than any number in the deck. The baseline methodology is covered in building the business case, and consumption-side forecasting in AI content budget planning. Current credit tiers are on pricing.

    FAQ

    What's a realistic ROI figure for AI-generated content?

    For pure substitution — same output, lower cost — most teams land in the range of a 40–70% reduction in cost per published asset once internal time is included on both sides. Capacity gains are usually larger and are the more defensible story, provided you can show the incremental assets did something.

    How do I prove AI content performs as well as produced content?

    Run a controlled split within one campaign: same audience, same budget, same window, one success metric chosen in advance. Expect rough parity per asset and a real advantage from having more variants to test, which is a portfolio effect rather than a per-asset one.

    Should time savings count as ROI?

    Only if the released hours were visibly reallocated to something with a name. Unallocated "time saved" is the fastest way to lose a finance reader, because it's the claim they've seen fail most often. Say where the hours went.

    How long before ROI shows up?

    Six to eight weeks. The first two weeks typically look worse than baseline while the team learns, which is normal and worth pre-announcing. Cost per asset usually crosses below baseline somewhere in week four or five.

    What if the numbers come out mediocre?

    Check keep rate and tier mix before concluding the program failed. Flat keep rate almost always means prompts and workflows haven't been standardized, and premium-tier drift inflates cost without improving output. Both are fixable in a fortnight.

    To get a clean baseline of your own, take one recurring asset type, run a full week through the AI video generator, and log human minutes and keep rate as you go — those two numbers do most of the work in every formula above.