Guides

    Claude is the word-choice brain; pin the renderer

    Claude ids via OpenRouter take image parts and win on scripts. They still do not burn captions or pick Kling unless you name Kling.

    Versely Team5 min read

    Claude ids via OpenRouter take image parts and win on scripts. They still do not burn captions or pick Kling unless you name Kling.

    The agent has two catalogs. Chat models plan, talk, and fill tool arguments. Generation models render the frame. Claude sits in the first catalog. It is a word-choice brain: brand voice, scripts, long captions as copy, a still used as a brief. It is not VEED, not Kling, not a caption burn. Mix the knobs and you will think "switching to Claude changed the look" when a router quietly picked a different renderer.

    Word choice is the job

    Picking the chat model behind the agent is the two-catalog rule. Chat ids live on GET /api/v1/agentic/chat-models. Generate ids live on the public catalog. Send chat_model as the id, not the marketing name. Claude ids are served via OpenRouter and, on the built-in registry, accept image parts — same vision side as Gemini 3.1 Pro, Grok 4.3, and Grok 4.6 (x-ai/grok-4.6). GLM 5.3 (z-ai/glm-5.3) is text-only; the image part 404s. Claude will not.

    If the job is "write the VO in our voice, then generate the clip," Claude is a fair pin. If the job is "make it look like last month's ad," the pin that matters is the renderer. Claude will route to a different family if you never named one. That look shift is not Claude's aesthetic. It is a generate tool filling models without a constraint.

    The default planner is not Claude. Treat /chat-models as live; the documented default is runpod/kimi-k2.7-code (Kimi K2.7 Code), text-only. GLM 5.3 is a custom OpenRouter id, not the default. Pin Claude when you want Claude, and mark it as a choice.

    Vision is allowed; rendering is not

    Attach a pack shot, a screenshot, a locked still, and a Claude id can see it. That is the still-as-brief path. It does not mean Claude "drew" the next frame. The MP4 still comes from whichever video model the generate tool was told to use.

    Video and audio attachments are pre-processed through Gemini for non-Gemini brains. You do not have to pick Gemini just to describe a clip. You do have to pick a vision chat model if the image is the brief. Claude is that pick when the brief is words-plus-a-still and the words matter more than a long tool-calling loop.

    On POST /api/v1/agentic/chat (and /chat/stream) send:

    • chat_model: the Claude id from /chat-models
    • chat_model_explicit: true when this is a choice, not a leftover default

    On the OpenRouter path, fast-eligible sub-agents (generation, slideshow, social, web editor) may stand in Claude Haiku 4.5 unless that flag is set. Workflow and movie planning keep the brain you picked. The default RunPod Kimi is not swapped. Omit the flag and a mechanical domain may run on Haiku while the picker still shows Claude. The generate model is untouched. The look is still the renderer.

    A lasting preference can store "always plan with this Claude id." It still does not pin Kling.

    Pin Kling in the same brief

    9 chat models to pin behind the Versely agent puts Claude in the word-choice slot: brand voice, scripts, long captions. Then name the pixel row. Kling 3 Turbo text-to-video if you have no plate and you accept invention. Seedance 2.5 if the still is the product and the job is still-to-motion. Do not write "Claude, pick the best video model" and then blame Anthropic for a family you did not lock.

    Prompt shape:

    Use this chat model for planning. Do not change the video model. Generate with Kling 3 Turbo, 5s, 9:16, locked-off bottle on wet stone. If you need allowed durations, call the schema tool, then generate. Do not substitute a different renderer.

    That message does two jobs. chat_model selects the planner. The named renderer selects the pixels. If you omit the second sentence, a newly eloquent brain may "help" by routing. You will credit Claude for a look it did not render.

    Chat tokens and generations bill on two meters. Changing brains changes the reasoner line. It does not change what a named Kling job costs. There is no free chat allowance.

    Captions are a burn, not a Claude paragraph

    Claude will write the lines. Add captions to a video burns them. add_veed_captions transcribes speech and typesets a preset; authored overlay is a different tool if nobody said the line out loud. Asking Claude to "put captions on it" without naming the caption tool is how you get a script in the chat pane and a silent file.

    Claude Code, Cursor, and Codex are a third door. Skills and MCP submit jobs through Versely. That host-tool Claude is not chat_model. Install skills on the laptop; pin the brain when you are in /agent.

    FAQ

    Will switching to Claude change my next clip's look?

    Only if the new brain picks a different generate model than the last one did. Pin the renderer in the brief (or in a saved workflow) if matching matters. Claude changes planning, tool use, and the text reply.

    Is Claude the default Versely chat model?

    No. Treat /chat-models as live. The documented default is runpod/kimi-k2.7-code. Claude is an OpenRouter-served vision id you send or save. GLM 5.3 is not the default either.

    Can Claude burn captions if I attach the cut?

    It can draft the words. Burning styled, timed captions is add_veed_captions (or an authored overlay tool). Name the editor job. A paragraph in the chat is not a burn-in.

    Do I pin Claude or Kling for a talking product ad?

    Both. Claude (or another vision brain) for the script and the still-as-brief. Kling, Seedance, or Veo — named — for the file. One id plans. The other renders.