App Store Screenshot Design From Generated Frames
App Store screenshots are a design job with a hard spec: up to 10 slots, three previews, one 1024x1024 icon. Build captioned frames that fit inside it.
An App Store listing is one of the few marketing surfaces where the spec sheet is the brief. There's no ambiguity to design around — Apple states exactly how many slots exist, what formats they accept, and what the one non-negotiable asset has to measure. The actual design job is what goes inside those fixed dimensions: which app moment earns a slot, and what caption over it does the work of a thirty-second pitch in the two seconds before a thumb keeps scrolling.
The spec is the brief
Three numbers govern the whole listing. Apple accepts one to ten screenshots per app, in .jpeg, .jpg or .png, with no alpha channel or transparency, at sizes that vary by device — a 6.9" iPhone display and an 11" iPad are different pixel dimensions entirely, and a 13" iPad listing is required if the app runs on iPad at all. Up to three app previews are allowed per supported device size and language, running 15 to 30 seconds each, across iOS, iPadOS, macOS, tvOS and visionOS — macOS and tvOS previews specifically must be landscape. And the app icon submitted to App Store Connect is a single 1024×1024 pixel master file, flat, square, and without transparency — Apple generates every smaller size an app actually displays from that one file.
Ten slots, three previews, one icon. That's the entire canvas — which means the design decision that actually matters is what earns a place in it, not how to fill space that doesn't exist.
The deliverable is a caption over a real frame, not a screenshot
"App Store screenshot" undersells what actually ships. A raw screen capture rarely converts on its own — the frames that do the work are a real captured moment from the app with a short, benefit-first caption laid over or beside it, in a consistent type treatment across the set. That's a design job with two real inputs: the app's actual UI (not a mockup — App Store guidelines expect what a screenshot shows to be genuinely representative of the app), and the caption copy that turns a static screen into a stated benefit. Ten slots is enough to walk a genuine narrative arc — hook, core feature, secondary feature, social proof, CTA — rather than ten screens of the same feature captioned six different ways.
Building the set from real captures
Versely's slideshow tooling maps onto this job directly, because it's built around exactly this shape: a fixed sequence of images, each with its own caption, in a consistent style. create_slideshow_from_uploaded_images takes the app's actual captured screens as image_urls — the order they're passed in becomes the slide order — and when text_style is set, the tool auto-writes a caption per image based on a prompt describing the app and stamps it onto each frame in one pass, rather than hand-placing text on every screen separately. This is real screenshots going in, not generated UI mockups going in: the tool is explicitly built for a user's existing images, which is the honest, guideline-compliant way to source the base frames for this specific deliverable.
Handling more than one device size
The dimension problem is real — a caption-and-frame layout tuned for a 6.9" iPhone doesn't drop cleanly into an 11" iPad canvas at a different aspect ratio without redesign. The practical approach is treating each device size as its own pass through the same caption script rather than one design stretched across sizes: same ten-slot narrative, same caption copy, re-composed against each device's actual dimensions rather than scaled. That's slower than a single universal export, but it's the difference between text that sits where it was designed to sit and text that clips against a different frame's edge because the canvas underneath it changed shape.
The icon is its own single-frame job
The 1024×1024 icon doesn't get a caption, a device frame, or a narrative slot — it's judged entirely on legibility at a few dozen pixels on a home screen, which is a different design problem from a screenshot's job of pitching a scrolling thumb. It's a single text-to-image generation with the aspect ratio locked to 1:1 and the resolution set high enough to hold detail down to the smallest size it'll actually render at, treated as its own deliverable rather than an afterthought cropped from a screenshot's hero frame.
A Versely walkthrough: one caption-over-frame set
For a five-screen narrative arc across a device's full ten-slot allowance:
- Capture the real frames. Pull five actual screens from the app covering hook, core feature, a secondary feature, social proof and a CTA screen — genuine UI, not a mockup.
- Build the captioned set.
create_slideshow_from_uploaded_imageswith the fiveimage_urlsin narrative order andtext_styleset, plus apromptdescribing the app and the benefit each screen should land — the tool auto-writes and stamps one caption per frame in a consistent style. - Review and adjust per slide.
replace_slideshow_slideorreorder_slideshow_slidesif one caption undersells its screen or the sequence reads better reordered — this is a design pass, not a one-shot generation. - Generate the icon separately. A dedicated
generate_imagescall, 1:1 aspect ratio, no caption or device frame — legibility at small size is the only test that matters here. - Repeat the caption script per device size. Re-run the composed set against each additional required size (iPad, and any others the app supports) rather than stretching one export across dimensions it wasn't designed for.
Typography carries more weight here than almost anywhere else
A caption that has to land in under two seconds, at a size a thumb is scrolling past, is one of the least forgiving typography jobs in any format — legibility at speed beats personality every time. Versely's font catalog is worth checking against that specific constraint before locking a caption style: a typeface that's distinctive at a poster's viewing distance can be the exact one that blurs into an average scroll speed at phone-screen size.
Checking the models before the set
/models is the place to check which image models are actually suited to this — legible small text and clean UI-adjacent framing aren't every model's strength, and Versely's text-to-image tool is the direct path to the icon generation specifically. For the broader question of which workflows fit a given app or audience beyond the screenshot set itself, /for breaks down capability by use case.
FAQ
How many screenshots should I actually use, out of the ten allowed? As many as the narrative genuinely needs, not the maximum by default. Ten screens of diminishing-return captions convert worse than five or six that each earn their slot — the ceiling is a limit, not a target.
Can app previews replace screenshots entirely? No — both exist on the listing simultaneously, capped independently at up to ten screenshots and up to three previews. They serve different scroll behaviours: a screenshot is read in a glance, a preview asks for a few seconds of attention.
Does the icon need to match the screenshot set's caption style? Visually, ideally yes — shared colour and type language across icon and screenshots reads as one coherent listing. Functionally, no — the icon carries no caption at all and is judged purely on recognisability at small size.
Is it against guidelines to use AI-generated captions or frame treatments on real screenshots? The captions and design treatment are marketing copy layered over the screenshot, not the app itself — what matters for compliance is that the underlying screen shown is genuinely representative of the app's real UI, not a fabricated interface.
Ten slots, three previews, one icon — the constraint is fixed, so the entire design job collapses into one question worth spending the real time on: which real moment from the app earns each slot, and what two seconds of caption makes a scrolling thumb stop on it.