Qwen3.5 397B sees the still; Qwen Image renders it
Vision chat id versus /models/qwen-image-3-text-to-image. Two catalogs. Pin both or the slide will be a different model than the planner.
Vision chat id versus /models/qwen-image-3-text-to-image. Two catalogs. Pin both or the slide will be a different model than the planner.
Qwen3.5 397B is a vision-capable brain on the built-in chat registry. It accepts image parts. It plans, talks, and calls tools. It does not mint the JPEG. Qwen Image 3 is the generate row: 4 credits, up to 2K, Chinese and English text rendering, prompt rewriting on by default. Same family name. Opposite jobs. Send the chat id as chat_model. Name the generate slug in the brief. Mix those knobs and you will think "Qwen" changed the look when a router picked Flux.
Two catalogs, one family name
Picking the chat model behind the agent is the rule. Chat models live on GET /api/v1/agentic/chat-models. Generation models live on GET /api/v1/agentic/models and on the public catalog. Swap the first and the agent plans differently. The file still comes from whichever row the generate tool was told to use.
Qwen3.5 397B sits on the vision side of the built-in roster with Gemini 3.1 Pro, the Claude ids, Grok 4.3, and OpenRouter Kimi K2.5 / K2.6 / K3. The GLM 5 family, DeepSeek V4 Pro, MiniMax M2.7, Qwen3 Max Thinking, and the RunPod Kimi ids — including the documented default runpod/kimi-k2.7-code — 404 on an image part. Qwen3 Max Thinking is not Qwen3.5 397B. Do not copy the name from one and send the other.
The generate slug is qwen-image-3-text-to-image. Provider Qwen. content_type image. requires_image false. Categories: text-to-image. Features named: text_to_image, text_rendering, multilingual, prompt_rewriting. Billing is per_megapixel, published as 4 credits. Max output 2K. Aspects include 1:1, 3:2, 2:3, 4:3, 3:4, 16:9, 9:16, 21:9. That row is not a chat brain. Do not send it as chat_model.
Treat /chat-models as live. Do not hard-code this post as the picker. GLM 5.3 (z-ai/glm-5.3) is a custom OpenRouter pin, not the default. Grok 4.6 (x-ai/grok-4.6) is the other custom vision pin. Qwen3.5 397B is the built-in Qwen vision id. Check the id field before you type a marketing string.
Vision chat is not a JPEG
Attach the approved still and pin Qwen3.5 397B when the pack shot is the brief. The planner can see the bottle, the label, the crop. Then it should still name the renderer. If the next file is a still, name Qwen Image 3 (or Flux, or Ideogram) on purpose. If the next file is motion, name Kling 3 Turbo or Seedance 2.5. Pinning the Qwen brain does not select the Qwen image row. A clever planner may "help" by routing to a different family. That is the look-shift bug.
Video and audio attachments are pre-processed through Gemini for non-Gemini brains. You do not have to pick Gemini just to describe a clip. You do have to pick a vision chat model if the image is the brief. Qwen3.5 397B is that pick when you want a Qwen brain rather than Gemini, Claude, or Grok 4.6.
Do not attach a still to the default Kimi and expect Qwen to have seen it. Kimi is text-only. The image part 404s. Switching the brain because the stills look wrong is the wrong fork. The stills are the renderer.
Pin both or the slide drifts
On POST /api/v1/agentic/chat (and /chat/stream) send the chat id from /chat-models. Send chat_model_explicit: true when the value is a choice. On the OpenRouter path, fast-eligible sub-agents (generation, slideshow, social, web editor) may stand in Claude Haiku 4.5 unless that flag is set. Workflow and movie planning keep the brain you picked. The default RunPod Kimi is not swapped for Haiku.
Qwen3.5 397B is built-in, not an OpenRouter custom slug. Still set the flag when you picked it. A lasting preference can store "plan with Qwen3.5 397B." It still does not pin Qwen Image 3.
Prompt shape:
Plan with Qwen3.5 397B. The attached still is the product. Do not change the image model. Generate the slide with Qwen Image 3, 2K, 9:16, lock the label. Do not substitute a different renderer.
If you omit the second sentence, the slide can land on another text-to-image row. You will blame Qwen3.5 for a look it did not render. The agent is the planner surface. The generate row is the pixels. Two doors.
Nine chat models to pin is the install list. Item 9 is this brain: vision, not the Qwen image generate row. Two catalogs.
What Qwen Image 3 actually bills
Four credits on the per-megapixel meter as published. 2K ceiling. Prompt rewriting on by default — the generate row may rewrite the prompt the planner wrote. If matching last month's slide matters, say so in the brief, and treat rewriting as a setting on the image job, not a personality of the chat brain.
Qwen Image 3 is multilingual text rendering, Chinese and English. That is why you pick it when the slide has glyphs. It is not why you pick Qwen3.5 397B. The brain reading a still and the row drawing letterforms are sequential jobs. Credits for the reasoner and credits for the render are two meters. Changing brains changes the first number. It does not change what a named 4-credit still costs.
Do not ask the chat brain to "just make the image." That is a generate tool call with a models argument. Name qwen-image-3-text-to-image or accept whatever the router picks. Honesty about which catalog you are turning is the whole post.
FAQ
Is Qwen3.5 397B the same as /models/qwen-image-3-text-to-image?
No. One is a vision chat id on /chat-models. The other is a 4-credit, 2K text-to-image generate row. Pin both if both jobs matter.
Is Qwen3.5 397B the default Versely planner?
No. Treat /chat-models as live. The documented default is runpod/kimi-k2.7-code. GLM 5.3 is not the default either.
Can I send the generate slug as chat_model?
No. chat_model takes a chat id. The generate slug belongs in the brief (or the tool's models argument). Sending the image slug as a brain will not draw a JPEG.
Will pinning Qwen3.5 397B make the next slide look like Qwen Image 3?
Only if the new brain picks that generate row. Name Qwen Image 3 in the brief if matching matters. Name a different stills row if it does not.