Models and architecture

    Multimodal model

    A multimodal model handles more than one kind of data — text, images, audio, video — inside a single system rather than bolting separate tools together.

    The useful consequence is that inputs can be mixed. A model that understands both pictures and text can be handed an image and asked a question about it, or given a reference and a written instruction and asked to reconcile them. That is a different capability from a pipeline where an image tool and a text tool pass files to each other.

    It is also why generation modes proliferate. Once a model can condition on several kinds of input, image-to-video, reference-to-video and audio-driven generation stop being separate products and become different combinations of the same inputs.

    Multimodal does not mean equally good at everything. A model can read images excellently and produce mediocre ones, or generate strong video and handle audio as an afterthought, so the label describes the interface rather than the quality.

    In practice

    • Mixed inputs — text plus image plus audio — are the defining feature.
    • Strength in one modality implies nothing about strength in another.
    • Most creative pipelines still chain specialists rather than relying on one model for everything.

    The mistake to avoid

    Assuming a model that understands images can also edit them. Reading and generating are different capabilities that happen to share a word.

    Related terms

    The all-in-one AI studio for creators. 60+ models for video, image, voice, music and lipsync in a single app.