AI video for museums and cultural institutions
Exhibition promos, collection storytelling and multilingual visitor content on a public-sector budget, plus where generated imagery has no business being.
The constraint at most museums is not the size of the marketing budget. It is the shape of it. Money arrives tied to a specific exhibition, on a procurement cycle that takes longer than the exhibition run, and it cannot be spent on anything the funding line did not name. Meanwhile the comms function is one person, sometimes half a person shared with development, and they are expected to produce a launch film, social cutdowns, gallery screen loops, and something in three languages for the visitor centre.
That mismatch is what makes generative video genuinely useful here, and it is also what makes it dangerous. A tool that produces plausible historical imagery in ninety seconds is a liability inside an institution whose entire authority rests on the difference between an artefact and a picture of one. So start with the line you will not cross, then build the production plan behind it.
Where generated imagery does not belong
This is the section to send to your curators before you send them anything else.
Never generate an artefact, a specimen, a document, or a site. Not for a promo, not for a thumbnail, not as a placeholder in a rough cut that might escape. The moment a synthetic object appears in institutional communications, every real object you have ever published becomes a question. Film the thing. If you cannot film the thing, use the thing's existing documentation photography, or use nothing.
Never generate a historical scene and let it read as documentation. Reconstruction is a legitimate interpretive form and museums have used illustration and diorama for a century. The difference is that a diorama announces itself. A photoreal generated street scene from 1890 does not. If you produce reconstruction, it needs to be visibly interpretive in style, labelled on screen, and labelled in the caption. The synthetic media disclosure conventions are the floor here, not the ceiling.
Never put a generated face on a named real person. First-person historical narration is compelling and it is also the fastest route to a story about your museum that you do not want. If you want a voice from the past, use an actor, credit the actor, and say the words are drawn from a specific document.
Check loan agreements before anything touches a model. Objects on loan usually carry reproduction terms negotiated with the lender, and those terms were not written with image models in mind. Do not put a loaned object's photography into a generative pipeline while you work out what the agreement means.
What generated material is genuinely good for, in this order: title and end cards, typography treatments, motion graphics and data animation, abstract texture and transition backgrounds built from your own identity palette, wayfinding and logistics animation, and explicitly stylised illustration where the style itself is the disclosure. That list covers most of the frames in a promo. The artefacts fill the rest, and they are real.
The three formats that carry the year
Everything a small comms team produces tends to collapse into three jobs. Build each one once as a repeatable structure rather than as a bespoke project.
| Job | What it is | Length | Where it runs |
|---|---|---|---|
| Exhibition promo | One idea, real objects, generated frame around them | 20 to 30s | Paid social, homepage, partner channels |
| Collection storytelling | One object, one story, no exhibition tie-in | 45 to 90s | Organic social, YouTube, newsletter |
| Visitor logistics | Opening hours, access, ticketing, what to expect | 15 to 25s | Website, screens, translated variants |
The exhibition promo is the one everybody budgets for and the one with the shortest useful life. The collection story is the one nobody budgets for and the one that still earns views two years later, because it is not attached to a date. If you are choosing where to put the hours, put them in the second row.
Collection storytelling also has the best content-to-cost ratio in the building. You already have the object, the curator, and the research. What you have historically lacked is the production layer around a two-minute curator interview: the title card, the lower thirds, the maps, the transitions, the end card with a booking link. That layer is exactly what generation is for, and it is reusable across every object you film afterwards.
Multilingual visitor content is the highest-return job
Most institutions serve an audience whose languages the comms team does not speak, and most of them handle it by producing everything in one language and printing a translated leaflet. Video localisation used to require a studio booking per language, which is why it did not happen.
Versely's dubbing tool takes a finished video and produces a translated version with the original speaker's voice characteristics carried across. Two engines sit behind it. The default handles audio or video, works on files up to thirty minutes, and supports trimming to a section. The second is video-only, produces lip-synced translation, caps out around eight minutes, and covers a narrower language list. For a curator interview, the first is usually right. For a piece where a presenter is on camera in close-up, the lip-synced route reads better.
The practical numbers, and they are three different numbers rather than one. The engines read 38 languages between them. 31 of those carry named voices you can actually pick from. 25 are dubbing targets. Chinese is the clearest example of why the distinction matters: it is a dubbing target with no named voices behind it. Check your language against the specific list you need before you promise it to a funder. The voice-over language pages show what each engine covers.
One operational constraint that catches people out: dubbing runs on Versely-hosted media. A YouTube link to your own video will not work as a source. Upload the master, dub from that.
Captions matter more than dubbing for gallery screens, because gallery screens play silent. Versely's subtitle path transcribes the speech and burns styled captions in, with a basic preset tier and a premium animated tier that costs double the basic. For institutional work the plain tiers are almost always correct. The caption tools cover the whole set, and the Classic caption family is the one that will not fight your house typography.
Running it on a fixed credit budget
Versely bills in credits, and the number for your exact model, length and resolution is shown before you confirm the generation. That is the property that makes this workable against a funding line: you can price a deliverable before you commit to it, which is a sentence most finance officers have never heard from a video supplier.
A realistic pattern for a small team:
- Film the real material first. Objects, curator, gallery. A phone on a tripod with a lavalier is fine for everything except the hero promo.
- Generate the frame elements once per exhibition. Title card, transition set, end card, lower-third backing. These get reused across every cutdown from that show.
- Assemble in the editor and iterate at preview resolution. The editor is timeline-based and re-renderable, and a preview pass renders at 480p for free with a short per-user cooldown between passes. Get the cut right there.
- Export once. The final export carries a single charge regardless of how many clips are on the timeline, so a fourteen-clip promo and a three-clip promo export the same way.
- Dub and caption from the finished master, not from the rough. Re-dubbing after a re-edit wastes the whole pass.
Two details worth knowing for institutional use. There are no watermarks on any plan, which matters when the output goes on a gallery screen or into a funder report. And the default frame rate is 25 fps, which is the right default for European broadcast and screen delivery but worth confirming against whatever your AV supplier specified for the gallery hardware.
If you want a per-job breakdown before planning a season, the credits page explains how a generation is priced, and the cost of dubbing into ten languages is the closest published example to a typical multilingual visitor brief. There is no free allowance to lean on, so the planning has to be real. The upside is that the planning is possible at all.
FAQ
Can we use generated video for an exhibition we have not installed yet?
For the space, no. For the idea, sometimes. A promo that runs before install usually cannot show the gallery, and the temptation is to generate one. Do not. Use the objects, the poster artwork, and typography. An abstract, clearly designed promo is honest. A generated gallery that does not match the real room produces complaints on opening weekend.
How should we label AI-assisted production?
A short line in the caption and, where any interpretive imagery appears, an on-screen label at the point it appears rather than only in the end credits. Institutions are held to a higher standard than brands here, and the labelling costs you nothing in performance. Keep the source files and a record of what was generated, because a journalist or a funder may eventually ask.
Is a cloned curator voice acceptable for the multilingual versions?
With written, revocable consent from that curator, and with the languages agreed up front. The thing to avoid is cloning a voice and then producing content that person never approved the wording of. Treat the clone as a delivery mechanism for scripts they have signed off, not as a standing licence.
What is the minimum viable setup for a one-person comms team?
One filming day per quarter, one reusable frame kit, and one editor draft you keep reopening. That gets you an exhibition promo, three collection stories and a logistics clip per quarter, plus translated variants of the logistics clip, which is more finished video than most institutions of that size currently ship in a year. If you want the sector-specific breakdown, the museums page covers the surfaces in more detail.