Comparisons

    Sign overlay vs a signed version

    A tiny interpreter inset is not automatically accessible. Compare overlay versus a signed-led cut, framing, contrast, and when to hire talent.

    Versely Team8 min read

    Burning a postage-stamp interpreter into the corner of a generated explainer does not make the video signed. It makes a second picture that is often too small to read, too low-contrast to parse, and cropped so the hands or the face fall out of frame. WCAG 2.2 Success Criterion 1.2.6 Sign Language (Prerecorded) is Level AAA. It requires sign language interpretation for prerecorded audio in synchronized media. It does not say "a small person in the corner is enough."

    WCAG's own definition is the part teams skip: sign languages are independent languages, unrelated to the spoken language of the same country. An interpreter is translating, not waving the English words. Captions remain the Level A/AA baseline. Sign is a different channel for people whose language is a sign language and who may not read captions at the rate a caption track presents them. That is the intent of 1.2.6, not a decoration pass.

    Two honest shapes exist. An overlay keeps the original picture dominant and composites a signer. A signed version makes the signer the picture, and treats the original as optional inset or omits it. Pick on purpose.

    Overlay versus signed-led, in production terms

    Technique G54 is the overlay: include the interpreter in the video stream. The technique is explicit about the cost. The interpreter cannot be enlarged without enlarging the whole image. If the stream is too small, the interpreter is indiscernible. Fullscreen becomes a requirement, not a nicety, unless the interpreter portion itself can be resized.

    Technique G81 is the better overlay when you control the player: a synchronized signer video in a separate viewport, or an overlay the player can scale. Most brand pages do not have that player. They have an MP4 with a burned-in inset. That file is G54, with G54's limits.

    A signed-led cut inverts the hierarchy. The signer is a medium shot, waist or mid-chest to above the head, with room for signs that travel. The original demo, if it still matters, sits in a second window that the signer can point to. For an announcement, a policy change, or anything whose information is in the language rather than in a product close-up, this is the cut that is actually readable on a phone.

    Question Overlay (signer inset) Signed-led cut
    Who is the picture? Original video Signer
    When it works Lecture, demo, slides that must stay large Announcement, explainer, anything language-first
    Phone risk Inset collapses to unreadable Signer still fills the frame
    Resize Cannot enlarge the signer without enlarging everything Signer is already the frame
    Contrast fight Signer competes with busy footage You control the backdrop
    Production Composite after both recordings exist Shoot or generate the original as a reference, not as the master

    Neither shape is "more AAA" by default. A huge, well-lit overlay on a desktop lecture can be usable. A signed-led cut with the hands cropped at the wrist is not. Judge the frame a viewer actually gets.

    Framing, contrast, and the face

    Sign is not hands alone. Non-manual markers (eyebrows, eye gaze, mouth, head tilt, shoulder) carry grammar. A question and a statement can share a handshape and differ on the face. If the inset is so small that the eyebrows are a smudge, the grammar is gone even if every manual sign is correct.

    Production rules that hold for a human interpreter, whether you overlay or lead with them:

    • Frame mid-chest to above the head, centered, with margin so a sign that extends off the torso does not clip. Cropping to a talking-head close-up, the way a talking-head model wants to frame a face, is the wrong crop. Sign needs the arms.
    • Light the face. A rim-lit silhouette with readable hands still drops the non-manuals.
    • Clothe against the background. Solid top, high contrast with the wall or the key, no busy pattern that camouflages handshape. The signer should not wear the same hue as the backdrop.
    • Keep both hands and the face in every shot. If you are compositing, do not place the inset over a lower-third, a caption, or a product that you then refuse to move. Crop safety applies: design the inset for the smallest playback size, not the desktop hero.

    The overlay a talking head agent path will composite a second clip onto a base video, including a corner placement and a black-background key. That is a UGC reaction tool. It is the right compositor only if the second clip is already a correctly framed, correctly interpreted sign performance. It will not invent one from a face photo.

    Gaze and micro-expression matter here the same way they matter on a speaking presenter; the gaze and micro-expression notes are about faces models already fail to hold. On a signer, that failure is grammatical, not cosmetic.

    When to hire talent instead of generating a signer

    A lipsync avatar makes a face speak. It matches a mouth to an audio track. That is the opposite of sign language interpretation. Prompting "ASL interpreter in the corner" through a talking-avatar generator will usually return a person gesturing, smiling, and mouthing English. Deaf viewers will clock it immediately. Shipping it as access is worse than shipping captions alone, because it claims a language it did not produce.

    Hire a qualified interpreter, or better, Deaf talent signing original content, when any of these is true:

    • The video carries instructions, legal language, health information, or anything a person might act on.
    • You need a specific sign language (ASL is not BSL is not Auslan) matched to the audience.
    • You cannot guarantee handshape, location, movement, orientation, and non-manuals at the playback size you actually ship.
    • You were about to generate a "signer" from a still of a hearing actor.

    Keep the overlay compositor for the last step: a real performance, keyed and sized. Keep generated talking heads for spoken delivery. Do not split the difference.

    Captions still ship. Caption readability is the AA track. Audio description is a different track again, for viewers who cannot see the picture. Sign does not replace either. If you cannot resource a human signer, do not ship a fake one. Ship captions, a transcript, and an honest absence of a signed version.

    A practical decision

    1. Is sign in scope? 1.2.6 is AAA. Many organizations stop at AA (captions plus audio description where needed). If sign is in scope, pick a language and an audience, not "a signer."
    2. Does the original picture have to stay dominant? If yes, overlay, and size the interpreter for the smallest player you support. If no, shoot a signed-led cut.
    3. Can the player scale the signer? If yes, prefer a separate signer file (G81). If no, burn in large, or do not overlay.
    4. Is the performance human? If no, stop. There is no generated substitute that currently holds grammar, regional variation, and trust at the same time. That limit is the subject of a companion piece; the production rule is already enough: human interpreter, or no signed version.

    The overlay is a layout. The signed version is a cut. Neither is a filter you apply to a talking-head export.

    FAQ

    Does a small interpreter inset meet WCAG 1.2.6?

    Only if a person fluent in that sign language can actually receive the interpretation, including non-manuals, at the size the video plays. G54 says a stream that is too small makes the interpreter indiscernible. A 160-pixel-wide inset on a mobile file fails that in practice even if an interpreter was on set.

    Can I overlay a generated avatar and call it signed?

    No. A lipsync model is solving mouth-to-speech, not language-to-language translation. An inset of that output is a picture of a person waving, not an interpretation.

    Should the signer appear on every social crop?

    If you cannot keep hands and face readable after the crop, do not overlay on that crop. Ship a signed-led vertical instead, or skip sign on that placement and keep captions. A clipped overlay is not a partial win.

    Is International Sign a shortcut for one file worldwide?

    No. International Sign is a contact variety used in some international settings, not a drop-in replacement for national sign languages. WCAG tells authors to choose the sign language of the primary audience, and to provide more than one when the audience is genuinely multiple. That is a hiring and scoping problem, not a prompt.