Guides

    Faith orgs: the building still and captions, not a synthetic imam

    Do not generate a religious presenter. Plates plus captions.

    Versely Team4 min read

    Do not generate a religious presenter. Plates plus captions.

    A mosque, a church, a gurdwara — the job on a feed is almost never "invent a person who appears to lead." People who know the community will see a face that is not theirs. People who do not will take the clip as clergy. That is the wrong asset class.

    The Versely churches row is the same discipline in another building: a real moment from the room beats an event flyer, and captions are what silent scrolling reads. Mosques keep a harder line on presenters. Do not synthesize the person at the front.

    A generated imam is the wrong asset

    Image-to-video and talking-head rows will happily give you a bearded man in a kufi if you ask. That is the model doing what talking-head rows are for. It is not permission.

    You do not need a synthetic speaker to announce jumu'ah time, an iftar, a food drive, or a roof appeal. You need the building people will walk into, and words they can read with the sound off.

    If you have real, consented footage of the room, use that. A clip of the hall reduces the uncertainty of a first visit. Generate a person only if that person exists, consented, and is not being passed off as clergy.

    Architecture is the plate

    Shoot or generate the still of the building, the mihrab wall, the courtyard, the door at dusk. Keep it a plate. If you need a little motion — clouds, a slow push, leaves — that is image-to-video on a locked still, not a new identity. The still is the contract. The prompt describes only what moves.

    Do not prompt "an imam delivering a sermon in a grand mosque" as the way to decorate a still you already have. You will get a presenter. You asked for one.

    Text-to-image is for a plate that does not exist yet — an illustration of the facade, a night exterior, a simple graphic of the dome. It is not for a face that will be read as a religious authority. If the words belong on the picture (a date, a one-line invitation), that is lettering on a still, then motion, not a talking composite.

    Burn the line, do not invent the speaker

    Most of this content is watched muted. The churches FAQ is blunt: caption even when the message is spoken clearly, because the silent autoplay is the first watch.

    Two different tools, because they are two different jobs:

    • Speech already on the file. Add captions to a video runs add_veed_captions: auto-transcribe, time to the voice, burn a preset. 165 language codes, including ar-SA. It transcribes; it does not translate, and it does not invent a speaker.
    • No one is talking. A building plate with a prayer-time line is not a transcript. Write the line and overlay it. Transcription has nothing to listen to.

    Do not ask a native-audio video model to "say the announcement." That is how you get a synthetic voice attached to a synthetic face. Put the announcement in type. Keep the plate as architecture.

    The test is ugly and useful: if the clip still works with the face cropped out, you did not need the face.

    FAQ

    Can we generate a generic "community member" instead of an imam?

    Only if you are willing to label it as generated, and only if that person is not styled as clergy. The safer default is no generated person at all. A courtyard still plus a captioned time does the announcement without impersonation.

    Do captions work for Arabic or bilingual recaps?

    add_veed_captions transcribes in the language actually spoken, including ar-SA among 165 codes. It will not translate a recitation into English for you. If you need two languages on screen, that is copy you write, not a transcript you hope the model invents.

    What if we only have a logo, not a photo of the building?

    Make a plate from the logo and a simple field. Do not invent a mosque interior you do not have rights to, and do not invent a person to stand in it. Motion on a logo still is enough for a story; a synthetic prayer leader is not.

    Should we use a talking-head model on a real volunteer?

    If the volunteer consented, is not presented as an imam, and the clip is clearly them, that is footage, not a generate. Captions still belong on it. The prohibition is the synthetic religious presenter, not every human on camera.