Guides

    Gemini 3.1 Pro is the brain that preprocesses everyone's clips

    Pick Gemini when the image is the brief. Video and audio attachments are pre-processed through Gemini for non-Gemini brains so you do not have to pick it just to describe a clip.

    Versely Team5 min read

    Pick Gemini when the image is the brief. Video and audio attachments are pre-processed through Gemini for non-Gemini brains so you do not have to pick it just to describe a clip.

    Gemini 3.1 Pro is the native Google entry on the agent chat list. It accepts image parts. That is the still-as-brief brain on the first-party roster, next to Claude ids, Grok 4.3, OpenRouter Kimi K2.5 / K2.6 / K3, and Qwen3.5 397B. It is not Veo 3.1. Veo is a generate row: native audio, 4K tier, 4 / 6 / 8 second clips. Mixing those two knobs is how “I picked Gemini” becomes a different MP4.

    Two catalogs, and Gemini sits on the chat one

    Picking the chat model behind the agent is the rule. Chat models live on GET /api/v1/agentic/chat-models. Generation models live on GET /api/v1/agentic/models and the public catalog. Swap the first and the agent plans, talks, and picks tools differently. The file still comes from whichever renderer the generate tool was told to use.

    Treat /chat-models as live. The documented default is runpod/kimi-k2.7-code (Kimi K2.7 Code), text-only. Gemini 3.1 Pro is not that default. GLM 5.3 (z-ai/glm-5.3) is not that default either — custom OpenRouter, text-only, image part 404s. Pin Gemini because the attachment is a still you need the planner to see, not because you dislike the stills. The stills are the renderer.

    Send chat_model as the id, not the marketing name. Send chat_model_explicit: true when the value is a choice. On the OpenRouter path, fast-eligible sub-agents (generation, slideshow, social, web editor) may stand in Claude Haiku 4.5 unless that flag is set. Workflow and movie planning keep the brain you picked. Chat tokens and generations bill on two meters. Changing to Gemini changes the reasoner line, not what a named Seedance 2.5 job costs.

    Stills 404 on text-only brains; clips do not have to

    Not every chat model accepts image parts. Text-only built-ins include the GLM 5 family, MiniMax M2.7, Qwen3 Max Thinking, the MiMo line, DeepSeek V4 Pro, and the RunPod-hosted Kimi ids (including the default). Send an image to those and the provider 404s on the image part.

    The practical split:

    • The image is the brief — pack shot, screenshot, frame to edit. Pick a vision-capable brain. Gemini 3.1 Pro is the native Google path. Grok 4.6 (x-ai/grok-4.6) is the OpenRouter vision pin if you want xAI instead. Claude ids also accept images. Do not “fix” a 404 by asking GLM to imagine the product.
    • The attachment is video or audio, and the selected brain is not Gemini. The server pre-processes that attachment through Gemini and injects an understanding into the prompt so a text-only brain can still act. You do not have to pick Gemini just to describe a clip.
    • No attachments, lots of tool use. The default Kimi is a reasonable start. Switch if the planner is dropping arguments, not because you wanted Veo.

    That preprocess is a short note, not a full analysis pass. Prefer analyze_video when you need reusable frames. Do not run it on every generate “because Gemini already looked.”

    Pick Gemini for the still; name the renderer anyway

    A prompt that keeps the knobs apart:

    Plan with Gemini 3.1 Pro. The attached still is the product. Do not change the video model. Image-to-video on Seedance 2.5, 9:16, slow orbit, no new label. If you need allowed durations, call the schema tool, then generate.

    chat_model selects the planner that can see the jpeg. The named renderer selects the pixels. If you omit the second sentence, a newly clever brain may “help” by routing to a different family. You will blame Gemini for a look it did not render. A lasting preference can store “always plan with Gemini 3.1 Pro.” It still does not pin Veo.

    Pinning Gemini 3.1 Pro in Versely chat does not select Veo 3.1. Name Veo when the job is Veo. The nine-brain install list is 9 chat models to pin behind the Versely agent. Gemini is item 7: native Google, vision, preprocess for everyone else’s clips.

    What Gemini 3.1 Pro is not

    It is not GLM 5.3. GLM cannot see the still. It is not Grok Imagine. Imagine is the generate family. It is not Gemini 3.1 Flash TTS — a 4-credit speech row, not this chat id.

    It is also not a reason to attach a clip and switch the planner to Gemini “so it understands the video.” The preprocess already ran for non-Gemini brains. Switching is extra reasoner spend unless you needed Gemini’s native vision on a still. If the brief is already words, you do not need Gemini. If the brief is a photo, you do. If the brief is a clip you only needed described, you probably do not.

    FAQ

    Do I have to pick Gemini 3.1 Pro to attach a video?

    No. Video and audio attachments are pre-processed through Gemini for non-Gemini brains and injected as understanding. Pick Gemini when the image is the brief, or when you want that brain for planning. Do not pick it only to describe a clip.

    Can I attach a pack shot to the default Kimi brain?

    Not as an image part. RunPod Kimi ids are text-only; the image part 404s. Use Gemini 3.1 Pro, Grok 4.6, Claude, or another vision chat id. Or describe the still in text and skip the attachment.

    Will pinning Gemini 3.1 Pro make the next clip look like Veo?

    Only if the new brain picks a Veo generate row. Pin Seedance 2.5 or Kling 3 Turbo in the brief if matching matters. Gemini is the planner, not the renderer.

    Is Gemini 3.1 Pro the default chat model?

    No. Treat /chat-models as live. The documented default is runpod/kimi-k2.7-code. Gemini is the native Google vision entry you send or save.