Workflows

    Google Demand Gen Creative With AI

    Google Demand Gen creative with AI: building the full asset matrix — 1:1, 4:5 and 16:9 images plus video — from one concept, without ten separate shoots.

    Versely Team8 min read

    Demand Gen campaigns fail for a boring reason more often than a strategic one: the asset group is half empty. Google's visual inventory spans Discover, Gmail, YouTube in-feed, and Shorts, and each surface wants a different shape. When you upload three square images and one 16:9 video, the system has almost nothing to assemble from, delivery concentrates in whichever placement your assets happen to fit, and the campaign underperforms in a way that looks like an audience problem but isn't.

    The requirement is uncomfortable if you produce creative traditionally. A single concept properly built out means landscape, square, and portrait images, plus video in both horizontal and vertical, plus enough variation in each that the system has real choices to make. That's a shoot with a shot list designed around aspect ratios — expensive, slow, and the reason most Demand Gen asset groups stay thin.

    Generated creative changes the arithmetic here more cleanly than in almost any other channel, because the problem is fundamentally combinatorial rather than creative. You have one idea and you need it in nine shapes.

    Multiple ad creative variants laid out across a designer's screen

    What a complete asset group looks like

    Google publishes specific requirements and they shift, so verify current specs in the interface before a build. The working shape as of mid-2026:

    Asset type Ratios to cover Practical minimum Where it surfaces
    Image 1.91:1 (landscape) 3–5 Discover, Gmail, in-feed
    Image 1:1 (square) 3–5 Discover, Gmail
    Image 4:5 (portrait) 2–3 Discover feed, mobile
    Video 16:9 (horizontal) 1–2 YouTube in-feed
    Video 9:16 (vertical) 1–2 Shorts
    Video 1:1 (square) 1 In-feed variants
    Headlines 5 All
    Descriptions 5 All

    That's roughly a dozen visual assets per concept. Note the shape of the requirement: it isn't twelve ideas, it's one idea rendered twelve ways. Understanding that distinction is what makes the build tractable — you're not brainstorming twelve times, you're producing systematically from one locked concept.

    Build the concept once, then render it wide

    The workflow that keeps this from becoming a slog:

    1. Lock the concept and the references. One angle, one product presentation, one visual mood. Gather two or three clean product reference images and, if there's a presenter, one reference face. Everything downstream generates against these so the whole asset group looks like it came from the same campaign — which matters, because Google assembles combinations you never see in preview.

    2. Generate the hero landscape image first. Get one 1.91:1 image you're genuinely happy with. This becomes the visual anchor.

    3. Reframe rather than re-prompt. For the square and portrait versions, extend the hero image rather than generating fresh ones. Outpainting preserves the exact product, lighting, and composition while filling the new frame — a far better result than three independent generations that agree on the brief but disagree on the details. One creative in every aspect ratio walks through this specifically.

    4. Produce the variety layer. Beyond the hero, Google wants genuinely different images to test — different settings, different framings, lifestyle versus product-only. Generate three or four distinct scenes against the same references, then reframe each into the ratios you need. This is where volume pays: five scenes × three ratios is fifteen images, and it's one session.

    5. Animate for the video slots. Image-to-video from your strongest stills is the cheapest route to usable video assets, and it guarantees the video matches the images in the same asset group. Generate 16:9 and 9:16 natively rather than cropping.

    6. Write the text assets against the visuals, not before them. Headlines that reference what's actually in the image assemble better.

    Getting text right in generated images

    Demand Gen images frequently carry a short line of copy, and this is where generated assets historically embarrassed people. Text rendering has improved substantially — models built for typography now handle short headlines reliably across multiple languages — but it still needs a check.

    Rules that keep you out of trouble:

    • Short strings only. Three to five words. Long sentences are where errors creep in.
    • Use a typography-capable model. Seedream 5.0 Pro handles in-image text notably better than general-purpose image models, and the text-to-image studio is where to run those generations.
    • Proofread every asset with text on it. Every single one. A near-miss character in an ad running at scale is a slow, expensive embarrassment.
    • Prefer platform text fields where possible. Headlines and descriptions supplied as text assets are cleaner, translatable, and testable. Reserve in-image copy for cases where it's genuinely part of the design.

    Model behavior varies enough here that it's worth generating one test headline before committing a whole batch to a model you haven't used for text.

    Feeding the system enough variety

    Demand Gen is a machine-assembled format, which changes how you think about "enough assets." You're not choosing the winning combination — you're defining the space the system searches. A narrow space produces narrow results.

    Practical guidance:

    • Vary the meaningful dimensions, not the trivial ones. Different settings, subjects, and framings teach the system something. Five color-graded versions of the same shot don't.
    • Include at least one lifestyle and one product-only image in every group. They win on different surfaces.
    • Don't mix concepts inside one asset group. If you have two genuinely different angles, that's two asset groups. Mixing them produces incoherent combinations.
    • Watch asset-level reporting and replace the consistently weak performers rather than rebuilding the group. Google will tell you which assets are pulling weight.

    The refresh cycle

    Demand Gen fatigues more slowly than paid social, but it does fatigue, and the refresh is easier than the initial build because your references and concept already exist.

    A workable rhythm: refresh two or three images and one video per asset group each month, keeping the top performers in place. Because everything generates against the same locked references, new assets slot in without a visual mismatch — which is the failure mode when you refresh by commissioning a new shoot six months after the original.

    Save the whole build as a reusable workflow, and monthly refresh becomes a twenty-minute job: swap the scene descriptions, run, reframe, upload. For a broader view of how a single production session stretches across placements and channels, repurposing one brand video into 30 assets covers the same principle at campaign scale.

    FAQ

    How many assets does a Demand Gen asset group actually need?

    Aim for full coverage of every ratio slot rather than a specific total — realistically that's around a dozen visual assets plus five headlines and five descriptions per concept. Thin asset groups restrict where the campaign can deliver, which caps performance regardless of how good the individual creatives are.

    Should I crop images to fill the different aspect ratios?

    No — extend them instead. Cropping a landscape image to portrait discards most of the composition and frequently cuts the product or subject. Outpainting keeps the original intact and generates the additional frame area, which produces a set that visibly belongs together.

    Can I reuse my Meta or TikTok creative in Demand Gen?

    Vertical video often transfers to the Shorts slot reasonably well. Feed-native social imagery usually doesn't transfer to Discover and Gmail placements, which sit in a more editorial context and reward cleaner, higher-craft visuals. Treat transfers as a starting point, not a shortcut.

    How do I keep a dozen assets looking like one campaign?

    Lock reference images and generate everything against them rather than prompting each asset from scratch. Since the system assembles combinations you never preview, visual consistency across the whole group is what prevents mismatched pairings from reaching a real viewer.

    Is generated imagery acceptable for Google ad policies?

    Generated creative is not prohibited as such, but the usual policy rules apply in full — no misleading claims, no unsubstantiated implied results, no depictions of real people without rights, and accurate representation of the product being sold. Where a visual could be read as documentary evidence of an outcome, make its illustrative nature clear.

    Build one concept into a complete asset group — hero image, reframed ratios, variety scenes, and matching video — in a single session, starting in the text-to-image studio.