VEED Fabric 1.0 Text: the mouth pass after the picture exists (720p, 40cr)
VEED Fabric 1.0 Text wants a plate and a track. It is not how you invent the scene.
VEED Fabric 1.0 Text wants a plate and a track. It is not how you invent the scene.
The catalog: turn a photo into a talking video from just a script — the voice is auto-generated to match the character. Content type lipsync. Categories text-to-lipsync and image-to-lipsync. Requires an image. 720p max. Qualities 480p and 720p. 40 credits. Catalog audio true. Features: Script-driven, Auto voice, Image to talking video. No duration enum on the card.
The scene is the photo. The words are the script. Fabric is the mouth pass. It will not design the kitchen.
Photo plus script, then 720p talking
VEED Fabric 1.0 Text needs a face that already exists as a still. Crop it like a talking shot: head and shoulders, room you approved, brand you can screenshot. Then write the script as spoken sentences, not as a cinematic paragraph. The model auto-generates a voice to match the character. You are not picking a separate ElevenLabs id on this card. You are not prompting a new location.
If the plate is wrong, stop. Generate or shoot the still on a text-to-image row. Do not ask Fabric to "make it look like a studio" in the script. The script is what the mouth says. The photo is the room.
The sibling VEED Fabric 1.0 is the other Fabric card — still 40 credits, still 720p, still lipsync. Text is the script-in, auto-voice path. If you already have a voice file, that is a different lipsync shape. Read the card in front of you.
AI lipsync is the tool. Lipsync is the family. Best talking-head model is the ranking that includes this text-to-lipsync / image-to-lipsync row. VEED's roster is avatars, Fabric, lipsync, and background removal — not a landscape generator.
Forty credits is the mouth pass
40 credits is the catalog figure. The matrix is resolution-based per second (SD 8, HD/4K-labelled 15) with a floor at 40 and a long ceiling. You are buying talking pixels at 720p, not a 4K scene invent. 480p exists for drafts. 720p is the max. There is no 4K Fabric Text.
No duration list means you do not type "8 seconds" into a dropdown that is not there. The script's length is the performance. Write a hook, not a podcast. A 400-word monologue on a still is how 40 credits become a long, expensive blink.
Auto voice is a feature, not a brand-voice lock. If the series needs the same clone every week, generate the track on TTS and use a lipsync path that takes audio. Fabric Text will match this character on this still. Next week's still may get a cousin.
Do not invent the scene here
Wrong jobs for this row:
- "A woman in a kitchen explains the product" with no photo. That is text-to-video or a stills row first.
- "Make the room darker and have her walk." That is I2V or an edit. Fabric moves a mouth on a plate.
- Stacking Fabric on a native-audio clip that already spoke. Two voices, one face.
Right job: approved still + script you would actually say + 720p talking output. Then caption in video editing. The mouth pass comes after the picture exists.
FAQ
Can I use VEED Fabric 1.0 Text without a photo?
No. It requires an image. Categories are text-to-lipsync and image-to-lipsync. The text is the script, not the scenery. Invent the plate on a stills model, then come back.
Does the auto-generated voice mean I should not write a script?
Write the script. Auto voice is how the words are spoken, matched to the character. A missing script is a missing track. A missing still is a missing scene.
Is 720p a preview of a 4K Fabric tier on this card?
No. Max output resolution is 720p, with 480p as the lower quality. Do not promise a 4K talking master from Fabric 1.0 Text.
When should I use ElevenLabs plus lipsync instead of Fabric Text?
When the voice has to be a locked clone or a Multilingual v2 read you already approved. Fabric Text auto-generates a voice to match the still. That is convenient. It is not a series bible.