Onboarding video for decentralized trials
Decentralized trials move procedures into the patient's kitchen. A multilingual onboarding build, and why every locale you add re-enters IRB review.
In a site-based trial, the first time a participant uses the device, a coordinator is standing next to them. In a decentralized trial, the first time a participant uses the device, they are at their kitchen table with a courier box, a quick-start card, and nobody to ask.
That relocation is the whole story. The protocol did not get simpler; the burden of executing it moved onto someone who has never done it and will not be observed doing it. Deviations follow — a dose logged outside the window, a sample collected at the wrong time of day, a wearable paired to the wrong account, an eDiary entry backfilled three days later. Onboarding video is one of the few interventions that scales against that, because it can be watched at the moment of the task rather than remembered from a visit three weeks earlier.
It is also, and this is the part marketing teams get wrong, a study document. Not a piece of content.
What the video covers, and what it must not
Onboarding video is post-consent. It is not recruitment material, and confusing the two is the fastest way to a finding. Recruitment video answers "am I eligible" and is subject to its own regime. FDA's information sheet on recruiting study subjects treats direct advertising as the start of the informed consent process, expects the IRB to review the final material rather than a draft, and rules out claims that the test article is safe or effective for the purpose under investigation or that imply a certainty of favourable outcome.
Onboarding video answers "how do I do this correctly." The content set is procedural:
- What arrives and when. The courier box, the contents, what to do if something is missing or damaged.
- Device unboxing and pairing. The actual device, the actual app screen, the actual sequence, and what a successful pairing looks like.
- The measurement itself. Body position, cuff placement, breath technique, number of attempts, what a failed reading looks like on screen.
- Drug storage and handling. Temperature range, what to do about excursions, where it must not go, how to record what was taken.
- The diary window. When it opens, when it closes, what to do about a missed entry — and that a missed entry recorded honestly is worth more than a backfilled one.
- Sample self-collection. Timing relative to dose and meals, labelling, the return path, the courier cut-off.
- Televisit preparation. What to have to hand, how to join, what happens if the connection drops.
- Who to call, and for what. Adverse events, device faults, logistics. Three different numbers, and participants conflate them constantly.
Two prohibitions run across all of it. Nothing may promote the investigational product — no efficacy language, no benefit framing, no "this will help you." And nothing may contradict, extend or soften the protocol. If the protocol says a window is two hours, the video does not say "roughly two hours."
Treat it like a protocol document, not an asset
The operational discipline that makes this work is version control borrowed from regulatory rather than from marketing.
The IRB approves the final video, including imagery. Not the script. Not a storyboard you intend to revise later. The thing that plays. So the review packet should include a shot-by-shot table with the on-screen text transcribed, because that is what gets read against the protocol.
A frame change is a document change. Reshooting one step because the app updated its icon is an amendment, not a tweak. Build the numbering in from day one — v1.0, v1.1, with a change log naming the protocol section each change touches.
Every version has an expiry and a retirement path. Sites and participants will have downloaded copies. If v1.0 shows a pairing flow that no longer exists, you need to know who has it and how it gets pulled.
Nothing procedural gets generated. This is the hard line. Generated material is fine for title cards, chapter transitions, wayfinding animation, ambient background and narration. It is not fine for a demonstration of a procedure. A model rendering hands performing a fingerstick will render it plausibly and wrongly, and you will have shipped a protocol deviation to every participant in the study. Film the real device, the real kit and the real screens.
Numbers, labels, dose text and window times go on as typed overlays for the same reason — what on-screen text actually holds up in generated video explains why the generated route fails at exactly the character level that matters here. Timed text overlays is the mechanism.
The build
A workable structure for a study with four to six procedures:
- Source the script from the protocol and the instructions for use. Not from a coordinator's summary. Cite the section number next to each line so the reviewer can check it without hunting.
- Shoot the procedural core. One camera, over-the-shoulder and top-down, the real kit, a real pair of hands, no music. Shoot each procedure as a discrete clip with clean heads and tails so it can be re-cut without re-shooting.
- Capture the app and portal screens. Screen recordings from a test account, never a real participant record.
- Build the connective tissue separately. Chapter cards, the "what to do if" branches, the contact card — the layers that change most often, kept in their own track. The explainer video generator handles the non-procedural sequences.
- Cut each procedure as a standalone short as well as a chapter of the master. Nobody rewatches a 12-minute film to find the sample-collection step.
- Assemble as one timeline with named swap slots. Saving the edit as a reusable draft means a locale or an amendment is a re-render, not a rebuild.
- Check the export against the approved packet before it goes anywhere. A pre-publish check pass catches the caption that covers the dose card.
Every locale re-enters review
This is the part that surprises teams coming from consumer marketing, where localisation is a throughput problem. Here it is a regulatory one.
A translated participant-facing document is a new document. It typically requires a certified translation, often a back-translation or an equivalent attestation, and submission to the reviewing IRB or ethics committee — and where a study runs across multiple countries, to each local committee under its own national requirements. The English master being approved does not carry the Spanish version. Budget review time per locale, not once.
That has a direct consequence for how you localise the video:
| Layer | Changes per locale | Re-approval implication |
|---|---|---|
| Procedural picture | No | Reusable across all locales |
| On-screen text and overlays | Yes | Must match the certified translation exactly |
| Subtitles | Yes | Must match the certified translation exactly |
| Narration audio | Yes | New audio script is a new document |
| Contact card | Yes | Local numbers, local hours, local escalation path |
| Device UI screens | Sometimes | If the app localises, re-capture; if not, subtitle the English UI |
Given that, burned-in subtitles keyed to the certified translation are usually the more defensible route than dubbing, because the text on screen is verifiably the approved text and a reviewer can read it directly off the frame. Burned-in captions also survive being downloaded and passed around, which sidecar files do not. If you do dub — and there are good accessibility arguments for participants with low literacy — the dubbing tool output is a new audio script that goes into the packet alongside the text, and the choice between dubbing and regenerating in the language is worth making deliberately rather than by default.
Do not lip-sync a filmed presenter into another language for this use. The claim being made by a synced mouth is that this clinician said these words, and they did not. Narration over procedural footage carries no such claim.
Keep re-renders cheap, because there will be many
Amendments happen. Devices get firmware updates. A site adds a country. The economics of this format depend entirely on whether a change costs a re-render or a re-shoot.
Three habits do most of the work. Keep the procedural footage in a library indexed by protocol section, so an amendment touching section 7.3 tells you exactly which clips to pull. Keep every word that could change — window times, contact numbers, version stamps — in overlay text rather than in the spoken track, so a change is a text edit and a fresh export rather than a recording session. And use the editor's 480p preview pass to check the assembled timeline before committing to a final export; the preview pass is free, carries a short per-user cooldown, and the final export is charged once regardless of clip count. Iterating on previews before paying for the export covers the loop.
Fix the frame rate at the start and leave it alone. Exports default to 25 fps, and mixing source rates into one timeline is the usual cause of the stutter that gets logged as a quality complaint in a review round.
FAQ
Does onboarding video need IRB approval if it only restates the protocol?
Ask your IRB rather than assuming. Materials given to enrolled participants are generally reviewed as study documents, and "it only restates the protocol" is a judgement the reviewer makes, not the producer. The cost of asking is a form. The cost of not asking is a finding on a document every participant has watched.
Can we use a generated presenter to deliver the instructions?
For narration over procedural footage, with disclosure, that is a reasonable position. For a figure presenting as a study clinician giving instruction, no — that manufactures a credential, and in a study context the credential is doing real work. The safest split is a plain narrator voice and no synthetic human in frame at all.
Is a separate video per country worth it, or one video with language tracks?
One picture, many text and audio tracks, is the right structure — but ship separate exports rather than one file with selectable tracks. Participant-facing distribution goes through portals and links, selectable tracks get missed, and a participant who opens the wrong language will not go looking for the menu. Localisation strategy covers the general case; the trial difference is that each export is separately approved and separately versioned.