Industry

    Limits of synthetic sign language avatars

    Generated sign is often grammatically wrong. What avatars miss, when they harm trust, and the decision rule: human interpreter or none.

    Versely Team7 min read

    A generated "ASL avatar" is usually a talking-head model doing extra homework. The mouth is matching English. The hands are improvising. The face is using expressions the way a hearing actor uses them, as emotion, not as grammar. Deaf viewers do not experience this as a rough draft of access. They experience it as a claim that a language was provided when it was not.

    That is not a taste argument. Signed languages are full languages. WCAG 2.2 defines them as combinations of hand and arm movements, facial expressions, and body positions, and it notes that they are independent of the spoken language of the same country. Word-for-sign substitution is not translation. An avatar that maps tokens to canned signs is doing the substitution.

    What current avatars miss

    The World Federation of the Deaf and the World Association of Sign Language Interpreters said this in a joint statement dated 14 March 2018, updated 14 April 2018. WFD and WASLI caution against signing avatars as a replacement for human signers. The reasons they list have not been retired by better rendering:

    1. Direct word-for-sign translations often do not exist. Equivalence needs lexicon, grammar, semantics, and discourse, not a dictionary lookup.
    2. Interpreters prepare. They account for who the audience is. An automated pass does not.
    3. Digitized corpora are incomplete, including regional and socio-linguistic variants, so the resource for generating signed statements is not there.
    4. Avatars are one-way, spoken or written into sign, not back.

    They single out live, complex, or high-stakes content: news, public emergency announcements, political announcements. They allow a narrow case for pre-recorded static customer information (a hotel check-in instruction, a queue sign) only if deaf people have advised on the signed sentences and no interaction or live signing is required.

    That is the bar. Most marketing video is not a hotel queue sign.

    Non-manuals are grammar. In American Sign Language, eyebrows, eye gaze, mouth patterns, head position, and shoulder shift mark questions, negation, relative clauses, and adverbial information. A signing model that animates a pleasant face while the hands move is dropping a channel. You cannot "add a smile" in post to recover a wh-question.

    Regional and community variation is not a skin. ASL is not British Sign Language. Black ASL is not a costume. Even inside one language, signs vary by region and by community. WFD's statement is explicit that dictionary work should document variation rather than pick one sign per word, and that WFD and WASLI do not support formal standardization of any sign language. An avatar trained on a single "correct" citation form will flatten that on purpose.

    Hands still break. Generated video already deforms hands when they grip a product and exposes itself on hands, teeth, and signage. Sign language encodes meaning in handshape, location, movement, and orientation. A fused finger is not an artifact you can ignore. It is a wrong phoneme.

    Lipsync is the wrong engine. Versely's talking-avatar path animates a face to speak provided audio, matching mouth movement to words. That is how lipsync models work: they are solving speech. They are not interpreters. Pointing that stack at "sign language" does not change the objective function. The talking-head models are the right tools for a spoken presenter. They are the wrong tools for a signed one.

    When a fake signer harms trust

    A missing signed version is a gap. A synthetic signer is a statement. The statement is: we knew this needed a language, and we generated a picture of one.

    That lands badly in three recurring situations.

    High-stakes information. WFD and WASLI already named news, emergencies, and political speech. Product-adjacent versions exist: a recall, a safety warning, a benefits explainer, a medical device how-to. If a person might act on the message, a grammatically wrong avatar is not a draft. It is a liability.

    Identity you do not have. Generating a "Deaf signer" from a hearing actor's still, or from a blended face that is not anyone, adds a likeness problem on top of a language problem. If the face is a real person, you need a likeness release that covers AI generation. If it is not, you are still presenting a person-like interpreter who cannot interpret.

    Sensitive subjects. A synthetic presenter on a topic the audience has reason to distrust is already a known failure mode for AI presenters and sensitive subjects. Sign language makes that sharper: you are not only choosing a face, you are claiming a community's language.

    The visual noise is part of the harm, not a separate QA nit. Jerky transitions between signs, floating hands, a torso that does not shift when the grammar requires it, a mouth that is visibly speaking English while the hands "sign": those are the tells. You will not catch them with a hearing review panel watching at 50 percent speed. You catch them by having fluent signers review, which is already most of the cost of hiring an interpreter.

    The decision rule

    Use this as a gate, not as a vibe:

    If you cannot put a qualified human interpreter, or Deaf talent signing original content, in front of the camera, do not ship a signed version.

    Ship what you can actually do:

    • Captions that meet SC 1.2.2, including speakers and meaningful non-speech sound.
    • A transcript that works as a document, not a cue dump.
    • Audio description where the picture carries information the soundtrack does not.

    SC 1.2.6 is Level AAA. Many accessibility policies stop at AA. That is a scoping choice you can defend. What you cannot defend is meeting 1.2.6 with a lipsync puppet.

    The narrow WFD/WASLI exception (pre-recorded, static, deaf-advised, no live interaction) is not a loophole for an ad campaign. A 30-second product film is not a train-station queue sign. If you think you are in that exception, you still involve deaf advisors on the actual signed sentences. You do not prompt an avatar and call the advisors later.

    When you do hire:

    • Specify the language (ASL, BSL, Auslan, the language your audience uses).
    • Frame for sign: mid-chest to above the head, hands always in, face lit.
    • Prefer a signed-led cut on mobile. An inset that dies at phone size is not interpretation.
    • Pay preparation time. Interpreters prepare; that is in the WFD/WASLI list for a reason.

    Generated video still has a job in the rest of the pipeline: the spoken cut, the B-roll, the stills. Leave the language that is not speech to people who have it.

    FAQ

    Have avatars improved enough since 2018 that the WFD/WASLI statement is obsolete?

    Image quality has improved. The statement's linguistic objections are about translation, corpora, preparation, and two-way communication, not about polygon count. Until those are solved in the language, a prettier puppet is still a puppet. The document has not been withdrawn.

    Can I use an avatar for internal drafts and replace it with a human for the ship?

    You can storyboard with a human on camera, or with a written gloss reviewed by a signer. Using a generated signer as the draft trains the rest of the team to accept wrong grammar as "close." It also risks leaking. If the draft must not be shippable, do not make it look like a signer.

    Does a disclaimer ("AI generated sign") fix the trust problem?

    It tells people you know. It does not make the language correct. For anyone who needs the information in that language, a labeled wrong translation is still wrong. For everyone else, it reads as a costume.

    What should a brand publish when the budget is captions but not an interpreter?

    Publish good captions and a readable transcript, and do not put a generated signer in the corner. An honest AA delivery beats a fake AAA badge.