Model Extra Parameters (extra_params)
Every model in model_config.toml accepts an extra_params field for passing provider-specific parameters in API calls. The most common use is controlling a model's thinking mode and reasoning intensity.
extra_params is not sent to the provider as-is. Before the request, the client splits and converts it by rule:
headers— passed as HTTP request headersquery— passed as URL query parametersbody— merged into the request body- Other plain keys — passed as extra request-body fields (the OpenAI SDK's
extra_body)
When client_type = "google", extra_params is not split by the rules above. Instead, the Gemini client filters the fields it supports and maps them to GenerateContentConfig.
Thinking and Non-thinking Modes
Many large models support a "thinking mode" — deep reasoning before answering, improving response quality on complex questions. MaiBot supports two API families, each configured differently:
- OpenAI-compatible API (
client_type = "openai"): DeepSeek, OpenAI, Alibaba Cloud Bailian, etc. - Gemini native API (
client_type = "google"): the Google Gemini family
WebUI Thinking Switch Adapts to Provider Templates
The thinking switch in the default-parameter area of the "Edit Model" dialog does not hard-code a parameter name. When you save a model, MaiBot matches the provider's base_url against the built-in provider templates, then writes the matching parameter into extra_params according to the template's thinking-format metadata. DeepSeek gets its own dedicated switch (including the Responses client and web search) — see DeepSeek Dedicated Switch. Every other provider falls into one of three write formats:
thinking_type— Zhipu, Kimi, MiniMax, Doubao: writesthinking.type. The "on" value differs per provider; templates that allow disabling writedisabledwhen offenable_thinking— Alibaba Cloud Bailian, SiliconFlow: writes the booleanenable_thinking = true/falsereasoning_effort— StepFun, OpenAI, xAI: reasoning models always think, so the switch cannot be turned off — only the effort level can be changed
When the Base URL matches no built-in template (a custom endpoint), no thinking switch is shown and you have to write extra_params by hand as described below.
Thinking effort (reasoning_effort) and thinking budget (thinking_budget) are two independent parameters: the effort selector appears only on templates that define effort levels, and the budget field appears only on templates that define a budget parameter. Leaving the budget empty means "use the provider default". Both only apply to OpenAI-compatible APIs; Gemini uses thinking_config.
The
model_identifiervalues below are placeholders — replace them with the actual model IDs you use.
thinking_type: Zhipu / Kimi / MiniMax / Doubao
All four write thinking.type, but the "on" value differs:
- Zhipu (GLM) — on value
enabled; effort parameterreasoning_effortwith levelsmax/xhigh/high/medium/low/minimal/none, defaultmax. Turning the switch on also writes the default level if no effort is set yet - Kimi (Moonshot) — on value
enabled, no effort levels - MiniMax (Hailuo AI) — on value
adaptive, letting the model decide; can be turned off - Doubao — on value
auto, letting the model decide; can be turned off; supports the thinking budgetthinking_budget
[[models]]
name = "zhipu-think"
model_identifier = "glm-5"
api_provider = "zhipu"
visual = false
extra_params = {thinking = {type = "enabled"}, reasoning_effort = "high"}[[models]]
name = "kimi-think"
model_identifier = "kimi-k3"
api_provider = "moonshot"
visual = false
extra_params = {thinking = {type = "enabled"}}[[models]]
name = "minimax-think"
model_identifier = "MiniMax-M3"
api_provider = "minimax"
visual = false
extra_params = {thinking = {type = "adaptive"}, reasoning_split = true}[[models]]
name = "doubao-think"
model_identifier = "doubao-pro"
api_provider = "doubao"
visual = false
extra_params = {thinking = {type = "auto"}, thinking_budget = 8192}Key points:
- Turning thinking off writes
thinking.type = "disabled"and also removes the effort parameter (meaningless once thinking is off). The budget parameter is kept - MiniMax chain-of-thought content is only returned in
reasoning_contentwhenreasoning_split = trueis set thinking_budgetonly accepts a non-negative integer; a negative number or a string is rejected by the switch as a configuration error
enable_thinking: Bailian / SiliconFlow
These two use a boolean switch, with no thinking object:
- Alibaba Cloud Bailian (Qwen) —
enable_thinking = true/false, plus the thinking budgetthinking_budget. Hybrid-thinking models also accept a/thinkor/no_thinkprompt suffix to switch temporarily - SiliconFlow —
enable_thinking = true/false, plus the thinking budgetthinking_budget
[[models]]
name = "qwen-think"
model_identifier = "qwen-plus"
api_provider = "dashscope"
visual = false
extra_params = {enable_thinking = true, thinking_budget = 8192}[[models]]
name = "qwen-nothink"
model_identifier = "qwen-plus"
api_provider = "dashscope"
visual = false
extra_params = {enable_thinking = false}[[models]]
name = "silicon-think"
model_identifier = "Qwen/Qwen3-32B"
api_provider = "siliconflow"
visual = false
extra_params = {enable_thinking = true}Key points:
enable_thinkingonly accepts a boolean; a string such as"true"is rejected with "enable_thinking can only be true or false"/thinkand/no_thinkare one-off prompt suffixes — they affect the current turn only and never change the model configuration
reasoning_effort: StepFun / OpenAI / xAI
These three templates have effort levels only and no switch: reasoning models always think, so the enable-thinking switch is greyed out and only the level can be changed. The allowed levels and defaults differ:
- StepFun —
low/medium/high, defaultmedium - OpenAI —
minimal/low/medium/high, defaultmedium - xAI (Grok) —
low/high, no default (the first level is used)
[[models]]
name = "openai-reason"
model_identifier = "gpt-5"
api_provider = "openai"
visual = false
extra_params = {reasoning_effort = "minimal"}[[models]]
name = "grok-reason"
model_identifier = "grok-4"
api_provider = "xai"
visual = false
extra_params = {reasoning_effort = "low"}[[models]]
name = "step-reason"
model_identifier = "step-3"
api_provider = "stepfun"
visual = false
extra_params = {reasoning_effort = "high"}Key points:
- An effort outside the allowed levels makes the WebUI report "reasoning_effort can only be …". When editing
model_config.tomlby hand, follow each provider's level list, otherwise the provider rejects the call - The WebUI only lists the levels configured in the template. OpenAI itself also accepts
noneandxhigh(see OpenAI-compatible APIs below); you can write them by hand, they are just not shown in the UI - Not being able to turn thinking off does not mean you cannot save tokens: drop to the lowest level (
minimalon OpenAI,lowon xAI) — that is the cheapest setting
DeepSeek Dedicated Switch
DeepSeek's two clients differ too much to fit the three general formats, so the model config page renders a dedicated block for it:
- Enable thinking — writes
thinking.typefor theopenaiclient, andreasoning.effortfor theopenai_responsesclient - Thinking effort —
low/high/max; greyed out while thinking is off - Enable web search — only available with the
openai_responsesclient; writes DeepSeek's nativeweb_searchtool and keeps any other tools already intools
[[models]]
name = "deepseek-think"
model_identifier = "deepseek-chat"
api_provider = "deepseek"
client_type = "openai"
extra_params = {thinking = {type = "enabled"}, reasoning_effort = "high"}[[models]]
name = "deepseek-web"
model_identifier = "deepseek-chat"
api_provider = "deepseek"
client_type = "openai_responses"
extra_params = {reasoning = {effort = "high"}, tools = [{type = "web_search"}]}Key points:
- Do not write
reasoningon theopenaiclient, and do not writethinkingorreasoning_efforton theopenai_responsesclient — mixing them is rejected with an "extra parameter conflict" error - The web-search tool may only be configured once, and configuring
web_searchon theopenaiclient is rejected outright; switch toopenai_responsesinstead - Full syntax: see Responses API below
Per-provider Limits and Pitfalls
These limits come from the providers themselves — violating them returns a 400 error. Check them before configuring:
- Kimi k2.7-code — only supports thinking enabled; writing
thinking.type = "disabled"fails - Kimi k3 series — does not accept the
thinkingparameter; onlyreasoning_effortcan tune the level, so do not writethinking - MiniMax M2.x — thinking cannot be turned off;
disabledonly works on M3 - Zhipu GLM-5.3 series — does not support turning thinking off, and the effort levels narrow to
low/high/max(defaultmax). GLM Coding Plan is a separate template and is not affected - Volcano Ark Coding Plan — plan credentials are not interchangeable with pay-as-you-go API keys; the console may show an account-specific Base URL that differs from the template, so follow the console
The matching extra_params writes:
[[models]]
name = "kimi-k27-code"
model_identifier = "kimi-k2.7-code"
api_provider = "moonshot"
visual = false
extra_params = {thinking = {type = "enabled"}}[[models]]
name = "kimi-k3"
model_identifier = "kimi-k3"
api_provider = "moonshot"
visual = false
extra_params = {reasoning_effort = "high"}[[models]]
name = "minimax-m3"
model_identifier = "MiniMax-M3"
api_provider = "minimax"
visual = false
extra_params = {thinking = {type = "disabled"}, reasoning_split = true}[[models]]
name = "glm-53"
model_identifier = "glm-5.3"
api_provider = "zhipu"
visual = false
extra_params = {thinking = {type = "enabled"}, reasoning_effort = "high"}Note
The four blocks above show hand-written extra_params. The WebUI switch only checks the template for whether disabling is allowed; per-model limits appear as small print below the switch (for example Kimi k2.7-code and MiniMax M2.x). Zhipu GLM-5.3 is a restriction the template applies per model identifier, so its switch is greyed out directly. Writing disabled in the source file for a model that does not support it makes saving report "the current model does not support turning thinking off".
New Plan-based Provider Templates
MaiBot 1.3.1 adds three subscription-plan provider templates. Their thinking switches match the corresponding pay-as-you-go templates:
- GLM Coding Plan —
https://open.bigmodel.cn/api/coding/paas/v4 - StepFun Step Plan —
https://api.stepfun.com/step_plan/v1 - Volcano Ark Coding Plan —
https://ark.cn-beijing.volces.com/api/coding/v3
[[api_providers]]
name = "zhipu-coding"
base_url = "https://open.bigmodel.cn/api/coding/paas/v4"
api_key = "your-glm-coding-plan-key"
client_type = "openai"[[api_providers]]
name = "stepfun-plan"
base_url = "https://api.stepfun.com/step_plan/v1"
api_key = "your-step-plan-key"
client_type = "openai"[[api_providers]]
name = "ark-coding"
base_url = "https://ark.cn-beijing.volces.com/api/coding/v3"
api_key = "your-ark-coding-plan-key"
client_type = "openai"Key points:
- Plan credentials are not interchangeable with pay-as-you-go keys — plan templates need the key issued by the plan subscription; filling in a pay-as-you-go API key fails authentication
- The Base URL may differ per account — the Volcano Ark Coding Plan console sometimes issues an account-specific address; use that address instead of copying the template
- Switch behaviour matches the corresponding pay-as-you-go template: GLM Coding Plan uses
thinking_type+reasoning_effort, Step Plan usesreasoning_effort, and Ark Coding Plan usesthinking_type(on valueauto) +thinking_budget
OpenAI-compatible APIs
The thinking object is a thinking-mode switch shared by several providers. DeepSeek, Kimi (Moonshot), and GLM (Zhipu) all use this format in exactly the same way. reasoning_effort is optional; omit it to use the default intensity. Some third-party platforms (e.g. Alibaba Cloud Bailian / DashScope) use the enable_thinking parameter format instead. The WebUI switch writes exactly the parameters described in this section, so hand-written configuration follows the same rules:
[[models]]
name = "deepseek-v4-flash-think"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
visual = false
extra_params = {thinking = {type = "enabled"}, reasoning_effort = "high"}[[models]]
name = "deepseek-v4-flash-nothink"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
visual = false
extra_params = {thinking = {type = "disabled"}}[[models]]
name = "deepseek-v4-flash-max"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
visual = false
extra_params = {thinking = {type = "enabled"}, reasoning_effort = "max"}[[models]]
name = "deepseek-v4-flash-think"
model_identifier = "deepseek-v4-flash"
api_provider = "dashscope"
visual = false
extra_params = {enable_thinking = true}[[models]]
name = "deepseek-v4-flash-nothink"
model_identifier = "deepseek-v4-flash"
api_provider = "dashscope"
visual = false
extra_params = {enable_thinking = false}Key points:
- DeepSeek V4's
reasoning_effortonly supports two valid levels:high(default) andmax(maximum reasoning).low/mediummap tohigh, andxhighmaps tomax - Compared with OpenAI: OpenAI's
reasoning_effortsupports 6 independent levels (none/minimal/low/medium(default)/high/xhigh), each effective on its own — unlike DeepSeek V4's 2 valid levels. Note thato1-minidoes not support this parameter - Multi-turn rule: if a thinking turn contains no tool calls, there is no need to send the thinking content back; if there are tool calls, it must be sent back
- Limitations: in thinking mode
temperatureandtop_pare silently ignored, andtool_choicecauses a 400 error - Third-party platforms (e.g. Alibaba Cloud Bailian / DashScope) control thinking mode with the
enable_thinkingboolean, which differs from the nativethinkingobject. Confirm which format your platform supports before configuring
Responses API
Some providers (e.g. DeepSeek v4 flash web search) use OpenAI's Responses protocol and require client_type = "openai_responses" in the provider configuration. The Responses API parameter format differs slightly from Chat Completions:
- Thinking mode: use
reasoning = {effort = "..."}instead ofthinking/reasoning_effort;effortacceptsnone/low/high/max - Web search: add the native
web_searchtool to thetoolslist to enable it
[[models]]
name = "deepseek-v4-flash-responses-think"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
client_type = "openai_responses"
extra_params = {reasoning = {effort = "high"}}[[models]]
name = "deepseek-v4-flash-responses-nothink"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
client_type = "openai_responses"
extra_params = {reasoning = {effort = "none"}}[[models]]
name = "deepseek-v4-flash-responses-web"
model_identifier = "deepseek-v4-flash"
api_provider = "deepseek"
client_type = "openai_responses"
extra_params = {reasoning = {effort = "high"}, tools = [{type = "web_search"}]}Key points:
- In the Responses client, do not write
thinkingorreasoning_effort; always usereasoning.effort, otherwise the request is rejected by validation reasoning.efforthas one more level than Chat Completions:none(fully disables thinking)- The
web_searchtool is passed viaextra_params.body.tools(fields in thebodygroup are merged into the request body together with other plain keys — see Custom HTTP Requests) - DeepSeek's Chat Completions endpoint does not support native web search; using web search requires the
openai_responsesclient - The Maisaka monitoring page and logs show a web-search summary for the turn (queries, actions, status, and source count)
Gemini Native API
When client_type = "google", extra_params is not processed by the OpenAI headers/query/body rules. The Gemini client filters the fields it supports and maps them to GenerateContentConfig.
Gemini 2.5 (thinking_budget)
The Gemini 2.5 family controls the thinking budget with thinking_budget (an integer):
[[models]]
name = "gemini-2.5-flash-think"
model_identifier = "gemini-2.5-flash"
api_provider = "google-gemini"
visual = true
client_type = "google"
extra_params = {thinking_config = {thinking_budget = 4096, include_thoughts = true}}[[models]]
name = "gemini-2.5-flash-nothink"
model_identifier = "gemini-2.5-flash"
api_provider = "google-gemini"
visual = true
client_type = "google"
extra_params = {thinking_config = {thinking_budget = 0}}[[models]]
name = "gemini-2.5-pro-think"
model_identifier = "gemini-2.5-pro"
api_provider = "google-gemini"
visual = true
client_type = "google"
extra_params = {thinking_config = {thinking_budget = -1, include_thoughts = true}}Key points:
thinking_budget:-1= automatic allocation,0= thinking disabled,N= a token budgetinclude_thoughts: whether the response includes the thinking process- Known issue: on Flash Preview 04-17, setting
thinking_budget = 0may fail
Gemini 3.0+ (thinking_level)
Gemini 3.0 and later control thinking intensity with thinking_level (an enum):
[[models]]
name = "gemini-3-flash-high"
model_identifier = "gemini-3-flash"
api_provider = "google-gemini"
visual = true
client_type = "google"
extra_params = {thinking_config = {thinking_level = "high", include_thoughts = true}}[[models]]
name = "gemini-3-flash-low"
model_identifier = "gemini-3-flash"
api_provider = "google-gemini"
visual = true
client_type = "google"
extra_params = {thinking_config = {thinking_level = "low", include_thoughts = true}}Key points:
thinking_levelaccepts:minimal,low,medium,high- Do not combine
thinking_budgetandthinking_level— it causes a 400 error - Multi-turn conversations need thought signatures to preserve context
Gemini Overview
- Gemini 2.5 — controls the thinking budget via
thinking_budget, range:-1(auto) /0(off) /N(budget); disable withbudget = 0. Budget and level cannot be combined - Gemini 3.0+ — controls the thinking level via
thinking_level, range:minimal/low/medium/high; disabled by leaving it unset or usingminimal. No token-level budget control
Gemini 2.5 controls intensity indirectly by token count (-1 = automatic), while Gemini 3.0+ sets the level directly with an enum value.
Google APIs are not directly accessible from mainland China; a proxy is required.
Custom HTTP Requests
extra_params supports three special keys for precise control over API requests:
headers— adds HTTP request headers, e.g.{headers = {"X-Custom" = "value"}}query— adds URL query parameters, e.g.{query = {"key" = "value"}}body— fields inside it go into the request body along with other plain keys; it exists only to group entries by purpose in the config
Note
body does not create a separate request channel. Fields inside body and all plain keys outside headers/query are merged and sent together in the request body.
For example:
[[models]]
name = "custom-model"
model_identifier = "custom-model-v1"
api_provider = "custom"
visual = false
extra_params = {
headers = {"X-API-Version" = "2024-06", "X-Priority" = "high"},
query = {version = "2024-01-01"},
body = {metadata = {source = "maibot"}},
enable_thinking = false
}The actual effect after the client splits it:
headers — HTTP request headers: X-API-Version: 2024-06, X-Priority: high
query — URL query parameters: ?version=2024-01-01
body fields + other plain keys — request body JSON: {"metadata": {"source": "maibot"}, "enable_thinking": false}
So extra_params = {enable_thinking = "false"} is equivalent to extra_params = {body = {enable_thinking = "false"}} — both send enable_thinking as a request-body JSON field to the provider, rather than a nested {"extra_params": {"enable_thinking": "false"}}.
Image Embedding Models
Image-memory retrieval needs the image_embedding task to use a model that "takes an image and returns a vector". The request shapes differ per provider, so with client_type = "openai" two template fields in extra_params adapt them:
image_embedding_input— the image input template,"{data_uri}"by default. Placeholders{base64},{data_uri}, and{mime_type}are replaced at request timeimage_embedding_body— the request-body template (an object). Itsinputkey overrides the input template above; all other keys are merged into the request body
When an official address is matched, MaiBot switches protocols automatically — no manual template needed:
- Alibaba Cloud Bailian, Volcano Ark (Doubao) — automatically switches to each provider's native multimodal embedding endpoint (Bailian
/api/v1/services/embeddings/multimodal-embedding/multimodal-embedding, Ark<Base URL>/embeddings/multimodal) and builds the body per the native protocol - SiliconFlow — still uses
/embeddings, with the{"image": "{data_uri}"}image input template filled in automatically - Other addresses — nothing is guessed; you must write
image_embedding_inputorimage_embedding_bodyyourself, otherwise the call fails
[[models]]
name = "vl-embed"
model_identifier = "vl-embed-v1"
api_provider = "custom-openai"
visual = false
extra_params = {image_embedding_input = "{data_uri}"}[[models]]
name = "qwen-vl-embed"
model_identifier = "multimodal-embedding-v1"
api_provider = "dashscope"
visual = false
extra_params = {dimensions = 1024}Key points:
- Under the Bailian/Ark native protocols,
modelandinput(plusencoding_formaton Ark) are reserved fields thatextra_paramsmust not override — writing them fails - Under the Bailian native protocol, the vector dimension can be written as top-level
dimensionsor asparameters.dimension; writing both with different values fails - Since 1.3.1, an
image_embeddingtask with no model configured fails immediately, telling you to specify an image-embedding-capable model underimage_embedding, instead of failing only at request time - The
embeddingandimage_embeddingtasks ignoreselection_strategyand always take the first available model inmodel_listorder — the vector space must stay consistent, since mixing embedding models corrupts the vector store
Advanced Auth Configuration
auth_header_name— Header auth name. DefaultAuthorizationauth_header_prefix— Header auth prefix. DefaultBearerauth_query_name— Query auth parameter name. Defaultapi_key
Other Advanced Parameters
Model-level Parameter Overrides
temperature— model-level temperature, overrides the task config. Optional, e.g.0.7max_tokens— model-level max tokens, overrides the task config. Optional, e.g.4096force_stream_mode— forces streaming output; settruewhen a model does not support non-streaming. Off by defaultextra_params— extra-parameter dictionary. Empty by default
Priority Rules
temperature and max_tokens can be written inside extra_params as model-level defaults, but prefer the standalone fields of the same name in the model config:
temperature = 0.7
max_tokens = 4096The intent is clearer, and it avoids confusion with same-named fields in the provider's request body.
When the same parameter exists in multiple places, the effective priority is:
- The value explicitly passed by the caller for this request
- The standalone field in the current model config (e.g.
temperature,max_tokens) - The same-named field in the current model's
extra_params - The default value in the current task config
API Provider Advanced Configuration
default_headers— default HTTP headers. Empty by defaultdefault_query— default query parameters. Empty by defaultorganization— OpenAI organization (optional). None by defaultproject— OpenAI project (optional). None by defaultmodel_list_endpoint— model-list endpoint. Default/modelsreasoning_parse_mode— reasoning-content parse mode. Defaultautotool_argument_parse_mode— tool-argument parse mode. Defaultauto
Runtime Configuration
timeout— request timeout. 60 seconds recommendedmax_retry— total attempts, first request included. 3 recommended, meaning up to 2 more tries after a failureretry_interval— retry interval. 5 seconds recommended
Quick Parameter Reference
OpenAI-compatible APIs
thinking— thinking-mode control, containstype.enabledfor Zhipu/Kimi,adaptivefor MiniMax,autofor Doubao,disabledto turn offenable_thinking— boolean thinking switch. Applies to Alibaba Cloud Bailian, SiliconFlowreasoning_effort— reasoning-intensity level. DeepSeek only low/high/max, OpenAI six levels, xAI only low/high, StepFun low/medium/high, Zhipu seven levels (defaultmax)reasoning— Responses API thinking control, containseffort(none/low/high/max). Applies to DeepSeek (Responses client)reasoning_split— MiniMax chain-of-thought split switch; withtruethe thinking content goes intoreasoning_contentthinking_budget— thinking budget (tokens). Applies to Doubao, Qwen/Bailian, SiliconFlowtools— native tool list, e.g.{type = "web_search"}enables web search. Applies to the DeepSeek Responses clientimage_embedding_input/image_embedding_body— image-embedding request templates. Applies to theimage_embeddingtaskheaders— custom HTTP request headers. Applies to allquery— custom URL query parameters. Applies to allbody— custom request-body fields. Applies to all
Gemini Native API
thinking_config— thinking configuration, containsthinking_budgetorthinking_level. Applies to the whole Gemini familythinking_budget— thinking budget (-1 auto / 0 off / N specified). Applies to Gemini 2.5thinking_level— thinking level (minimal/low/medium/high). Applies to Gemini 3.0+include_thoughts— whether the response includes the thinking process. Applies to the whole Gemini family
Parameters are passed to the LLM API verbatim — make sure they match your provider's documentation, otherwise calls may fail.
Verification and Troubleshooting
After configuring, run a capability test with the "Test model" button in the model list on the model config page: a thinking model returns text on success, and the test details show the reasoning content. If no reasoning content appears, check the thinking parameters in extra_params first.
- Saving reports "the current model does not support turning thinking off" → the template tightened the disable capability for that model identifier (typically Zhipu GLM-5.3). Kimi k2.7-code and MiniMax M2.x are per-model limits whose switch is not greyed out, so avoid writing
disabledwhen configuring by hand - Saving reports "reasoning_effort can only be …" → the level is outside that provider's range; fix it against the level lists above (xAI only has
low/high) - Saving reports "enable_thinking can only be true or false" → you wrote a string; change it back to a boolean
- A call returns 400 and the payload mentions
thinkingorreasoning_effort→ the client type and the parameter are mixed up: DeepSeek withopenai_responsesonly acceptsreasoning.effort, while DeepSeek withopenaionly acceptsthinking+reasoning_effort - Image-memory retrieval stays unavailable → confirm
[model_task_config.image_embedding]points to a model that accepts image input; for addresses other than Bailian, SiliconFlow, or Ark you must writeimage_embedding_inputorimage_embedding_bodyyourself
For details, rely on each provider's official documentation: DeepSeek thinking mode, OpenAI reasoning guide, Google Gemini thinking config, Alibaba Cloud Bailian, Kimi thinking mode, GLM thinking mode.