Skip to content

Model inference pricing

Tiered pricing rules

Some Model Studio models use tiered pricing. The total number of input tokens in a single request determines the unit price. All tokens in the request are then billed at the unit price for that tier.

For example, a model may have two pricing tiers: 0 < tokens ≤ 32K and 32K < tokens ≤ 128K. If a request contains 100K tokens, it falls into the second pricing tier (32K < 100K ≤ 128K). In this case, all 100K tokens are billed at the unit price for that tier.

Text generation: Qwen

Qwen-Max

You are charged for input tokens and output tokens.

If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. If the model supports context cache, only input tokens receive a discount. These two discounts cannot apply simultaneously.

International

With the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in your selected region. Supported region: Singapore.

Valid for 90 days after you activate Model Studio

Model nameModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**Free quota (Note)
qwen3.6-max-preview ** Discount for context cachingthinking and non-thinking0 < tokens ≤ 128K$1.3$7.81 million tokens for input and output eachValid for 90 days after you activate Model Studio
128K < tokens ≤ 256K$2$12
qwen3-max Currently equivalent to qwen3-max-2026-01-23 Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15
qwen3-max-2026-01-23thinking and non-thinking0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15
qwen3-max-2025-09-23non-thinking only0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15
qwen3-max-preview Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15
More models
Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
qwen-max ** Currently equivalent to qwen-max-2025-01-25 50% discount for batch inferencenon-thinking onlyno tiered pricing$1.6$6.41 million tokens for input and output each Valid for 90 days after you activate Model Studio
qwen-max-latestnon-thinking onlyno tiered pricing$1.6$6.4
qwen-max-2025-01-25non-thinking onlyno tiered pricing$1.6$6.4

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 Discount for context cachingnon-thinking only0 < tokens ≤ 32K$0.359$1.434
32K < tokens ≤ 128K$0.574$2.294
128K < tokens ≤ 256K$1.004$4.014
qwen3-max-2025-09-23non-thinking only0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.434$5.735
128K < tokens ≤ 256K$2.151$8.602
qwen3-max-preview Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.434$5.735
128K < tokens ≤ 256K$2.151$8.602

Chinese mainland

With the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3.6-max-preview ** Discount for context cachingthinking and non-thinking0 < tokens ≤ 128K$1.238$7.426
128K < tokens ≤ 256K$2.063$12.377
qwen3-max Currently equivalent to qwen3-max-2026-01-23 50% discount for batch inference Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$0.359$1.434
32K < tokens ≤ 128K$0.574$2.294
128K < tokens ≤ 256K$1.004$4.014
qwen3-max-2026-01-23thinking and non-thinking0 < tokens ≤ 32K$0.359$1.434
32K < tokens ≤ 128K$0.574$2.294
128K < tokens ≤ 256K$1.004$4.014
qwen3-max-2025-09-23non-thinking only0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.434$5.735
128K < tokens ≤ 256K$2.151$8.602
qwen3-max-preview Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.434$5.735
128K < tokens ≤ 256K$2.151$8.602
More models
Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens)
qwen-max ** Currently equivalent to qwen-max-2024-09-19non-thinking onlyno tiered pricing$0.345$1.377
qwen-max-latestnon-thinking onlyno tiered pricing$0.345$1.377
qwen-max-2025-01-25non-thinking onlyno tiered pricing$0.345$1.377
qwen-max-2024-09-19non-thinking onlyno tiered pricing$2.868$8.602

China (Hong Kong)

With the China (Hong Kong) deployment scope, both model inference compute resources and your static data are located in the China (Hong Kong) region. Note

Models in the China (Hong Kong) deployment scope do not offer a free quota.

Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 Discount for context cachingthinking and non-thinking0Model name**ModeInput tokens
---------------
qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 50% discount for batch inference Discount for context cachingthinking and non-thinking0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15
qwen3-max-2026-01-23thinking and non-thinking0 < tokens ≤ 32K$1.2$6
32K < tokens ≤ 128K$2.4$12
128K < tokens ≤ 256K$3$15

Qwen-Plus

You are charged for input tokens and output tokens.

International

When the International deployment scope is selected, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the Singapore region.

Model**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)Free quota** (note)**
Non-thinking modeThinking mode (CoT + response)
qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-020Model name**Input token rangeInput priceOutput price
---------------
Non-thinking modeThinking mode
qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-020 < token ≤ 256K$0.276$1.651$1.651
256K < token ≤ 1M$1.101$6.602$6.602
qwen3.6-plus-2026-04-020 < token ≤ 256K$0.276$1.651$1.651
256K < token ≤ 1M$1.101$6.602$6.602
qwen3.5-plus Currently equivalent to qwen3.5-plus-2026-02-150 < token ≤ 128K$0.115$0.688$0.688
128K < token ≤ 256K$0.287$1.72$1.72
256K < token ≤ 1M$0.573$3.44$3.44
qwen3.5-plus-2026-02-150 < token ≤ 128K$0.115$0.688$0.688
128K < token ≤ 256K$0.287$1.72$1.72
256K < token ≤ 1M$0.573$3.44$3.44
qwen-plus Currently equivalent to qwen-plus-2025-12-010 < token ≤ 128K$0.115$0.287$1.147
128K < token ≤ 256K$0.345$2.868$3.441
256K < token ≤ 1M$0.689$6.881$9.175
qwen-plus-2025-12-010 < token ≤ 128K$0.115$0.287$1.147
128K < token ≤ 256K$0.345$2.868$3.441
256K < token ≤ 1M$0.689$6.881$9.175
qwen-plus-2025-09-110 < token ≤ 128K$0.115$0.287$1.147
128K < token ≤ 256K$0.345$2.868$3.441
256K < token ≤ 1M$0.689$6.881$9.175
qwen-plus-2025-07-280 < token ≤ 128K$0.115$0.287$1.147
128K < token ≤ 256K$0.345$2.868$3.441
256K < token ≤ 1M$0.689$6.881$9.175

US

If you select the US deployment scope, model inference compute resources are located in the US. Static data is stored in your selected region. The supported region is US (Virginia). Note

Models in the US deployment scope do not offer a free quota.

Model**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode (CoT + response)
qwen-plus-us0 < token ≤ 256K$0.4$1.2$4
256K < token ≤ 1M$1.2$3.6$12
qwen-plus-2025-12-01-us0 < token ≤ 256K$0.4$1.2$4
256K < token ≤ 1M$1.2$3.6$12

Chinese mainland

Selecting the Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode (chain of thought and response)
qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-020Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
------------
qwen-plus-2025-01-25no tiered pricing$0.115$0.287
qwen-plus-2025-01-12no tiered pricing$0.115$0.287
qwen-plus-2024-12-20no tiered pricing$0.115$0.287

China (Hong Kong)

Selecting the China (Hong Kong) deployment scope restricts model inference compute resources to China (Hong Kong). Static data is also stored in this region. Note

Models in the China (Hong Kong) deployment scope do not offer a free quota.

ModelInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode
qwen-plus ** Currently equivalent to qwen-plus-2025-12-010 < tokens ≤ 256K$0.4$1.2$4
256K < tokens ≤ 1M$1.2$3.6$12
qwen-plus-2025-12-010 < tokens ≤ 256K$0.4$1.2$4
256K < tokens ≤ 1M$1.2$3.6$12

European Union

For deployments in the European Union, compute resources for model inference remain within the European Union. Static data is stored in your selected region. Supported region: Germany (Frankfurt). Note

Models in the EU deployment scope do not offer a free quota.

Model**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinkingThinking (Chain of Thought (CoT) + response)
qwen-plus ** Currently equivalent to qwen-plus-2025-12-010 < token ≤ 256K$0.4$1.2$4
256K < token ≤ 1M$1.2$3.6$12
qwen-plus-2025-12-010 < token ≤ 256K$0.4$1.2$4
256K < token ≤ 1M$1.2$3.6$12

Qwen-Flash

You are charged for input tokens and output tokens.

If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. If the model supports context cache, only input tokens receive a discount. These two discounts cannot apply simultaneously.

International

When you select the International deployment scope, inference computation is dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is supported in the Singapore region.

Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-16 50% off for batch inference Discount for context caching0 < tokens ≤ 256K$0.25$1.51 million input tokens and 1 million output tokens Valid for 90 days after you activate Model Studio.
256K < tokens ≤ 1M$1$4
qwen3.6-flash-2026-04-160 < tokens ≤ 256K$0.25$1.5
256K < tokens ≤ 1M$1$4
qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-23 50% off for batch inference Discount for context caching0 < tokens ≤ 1M$0.1$0.4
qwen3.5-flash-2026-02-230 < tokens ≤ 1M$0.1$0.4
qwen-flash Currently equivalent to qwen-flash-2025-07-28 50% off for batch inference Discount for context caching0 < tokens ≤ 256K$0.05$0.4
256K < tokens ≤ 1M$0.25$2
qwen-flash-2025-07-280 < tokens ≤ 256K$0.05$0.4
256K < tokens ≤ 1M$0.25$2

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-160 < tokens ≤ 256K$0.165$0.99
256K < tokens ≤ 1M$0.66$3.961
qwen3.6-flash-2026-04-160 < tokens ≤ 256K$0.165$0.99
256K < tokens ≤ 1M$0.66$3.961
qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-230 < tokens ≤ 128K$0.029$0.287
128K < tokens ≤ 256K$0.115$1.147
256K < tokens ≤ 1M$0.172$1.72
qwen3.5-flash-2026-02-230 < tokens ≤ 128K$0.029$0.287
128K < tokens ≤ 256K$0.115$1.147
256K < tokens ≤ 1M$0.172$1.72
qwen-flash Currently equivalent to qwen-flash-2025-07-28 Discount for context caching0 < tokens ≤ 128K$0.022$0.216
128K < tokens ≤ 256K$0.087$0.861
256K < tokens ≤ 1M$0.173$1.721
qwen-flash-2025-07-280 < tokens ≤ 128K$0.022$0.216
128K < tokens ≤ 256K$0.087$0.861
256K < tokens ≤ 1M$0.173$1.721

US

When you select the US deployment scope, inference computation is restricted to the US. Static data is stored in your selected region. This deployment scope is supported in the US (Virginia) region. Note

Models in the US deployment scope do not offer a free quota.

Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
qwen-flash-us0 < tokens ≤ 256K$0.05$0.4
256K < tokens ≤ 1M$0.25$2
qwen-flash-2025-07-28-us0 < tokens ≤ 256K$0.05$0.4
256K < tokens ≤ 1M$0.25$2

Chinese mainland

When you select the Chinese mainland deployment scope, inference computation is restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-16 50% off for batch inference Discount for context caching0 < tokens ≤ 256K$0.165$0.99
256K < tokens ≤ 1M$0.66$3.961
qwen3.6-flash-2026-04-160 < tokens ≤ 256K$0.165$0.99
256K < tokens ≤ 1M$0.66$3.961
qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-230 < tokens ≤ 128K$0.029$0.287
128K < tokens ≤ 256K$0.115$1.147
256K < tokens ≤ 1M$0.172$1.72
qwen3.5-flash-2026-02-230 < tokens ≤ 128K$0.029$0.287
128K < tokens ≤ 256K$0.115$1.147
256K < tokens ≤ 1M$0.172$1.72
qwen-flash Currently equivalent to qwen-flash-2025-07-28 Discount for context caching0 < tokens ≤ 128K$0.022$0.216
128K < tokens ≤ 256K$0.087$0.861
256K < tokens ≤ 1M$0.173$1.721
qwen-flash-2025-07-280 < tokens ≤ 128K$0.022$0.216
128K < tokens ≤ 256K$0.087$0.861
256K < tokens ≤ 1M$0.173$1.721

China (Hong Kong)

When you select the China (Hong Kong) deployment scope, inference computation is restricted to China (Hong Kong). Static data is stored in your selected region. This deployment scope is supported in the China (Hong Kong) region. Note

Models in the China (Hong Kong) deployment scope do not offer a free quota.

Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
qwen3.5-flash ** Currently equivalent to qwen3.5-flash-2026-02-23 Discount for context caching0 < tokens ≤ 1M$0.1$0.4
qwen3.5-flash-2026-02-230 < tokens ≤ 1M$0.1$0.4

European Union

When you select the European Union deployment scope, inference computation is restricted to the European Union. Static data is stored in your selected region. This deployment scope is supported in the Germany (Frankfurt) region. Note

Models in the EU deployment scope do not offer a free quota.

Model name**Input token rangeInput price (per 1M tokens)Output price (per 1M tokens)
qwen3.5-flash ** Currently equivalent to qwen3.5-flash-2026-02-23 Discount for context caching0 < tokens ≤ 1M$0.1$0.4
qwen3.5-flash-2026-02-230 < tokens ≤ 1M$0.1$0.4

Qwen-Turbo

Note

Qwen-Turbo will no longer be updated. We recommend using Qwen-Flash instead. You are charged for input tokens and output tokens.

If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price.

International

In the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

Model**Input price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
Non-thinking modeThinking mode (chain of thought + response)
qwen-turbo ** Currently equivalent to qwen-turbo-2025-04-28 50% discount for batch inference$0.05$0.2$0.51 million tokens each Valid for 90 days after activating Model Studio
qwen-turbo-latest$0.05$0.2$0.5
qwen-turbo-2025-04-28$0.05$0.2$0.5
More models
Model**Input price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
qwen-turbo-2024-11-01$0.05$0.21 million tokens each Valid for 90 days after activating Model Studio

Chinese mainland

In the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode (chain of thought + response)
qwen-turbo ** Currently equivalent to qwen-turbo-2025-04-28$0.044$0.087$0.431
qwen-turbo-latest$0.044$0.087$0.431
qwen-turbo-2025-07-15$0.044$0.087$0.431
qwen-turbo-2025-04-28$0.044$0.087$0.431

QwQ

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.

Model**Input price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
qwq-plus ** Currently equivalent to qwq-plus-2025-03-05$0.8$2.41 million tokens Valid for 90 days after you activate Model Studio

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are confined to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Input price (per 1M tokens)Output price (per 1M tokens)
qwq-plus ** Currently equivalent to qwq-plus-2025-03-05$0.230$0.574
qwq-plus-latest$0.230$0.574
qwq-plus-2025-03-05$0.230$0.574

Qwen-Long

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

Model**Input price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
** **qwen-long-latest$0.072$0.287No free quota
qwen-long-2025-01-25$0.072$0.287

Qwen-Omni

Billing is based on the number of input and output tokens. For token calculation rules for different modalities, see Billing and Throttling.

International

If you select the International deployment scope, model inference resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.

ModelInput priceOutput priceFree quota (Note)
Text/image/videoAudioText ** For multimodal inputText+audio** ** Audio billed only
qwen3.5-omni-plus Currently equivalent to qwen3.5-omni-plus-2026-03-15$1.4$11$8.3$441 million tokensThis quota expires 90 days after you activate Model Studio.
qwen3.5-omni-plus-2026-03-15$1.4$11$8.3$44
qwen3.5-omni-flash Currently equivalent to qwen3.5-omni-flash-2026-03-15$0.4$3$2.2$11.9
qwen3.5-omni-flash-2026-03-15$0.4$3$2.2$11.9
More models
Model**ModeInput priceOutput priceFree quota (Note)
TextAudioImage/videoText ** For text-only inputText** ** For multimodal inputText+audio** ** Audio billed only
qwen3-omni-flash Currently equivalent to qwen3-omni-flash-2025-12-01non-thinking and thinking modes$0.43$3.81$0.78$1.66$3.06$15.111 million tokens (modality-agnostic)This quota expires 90 days after you activate Model Studio.
qwen3-omni-flash-2025-12-01non-thinking and thinking modes$0.43$3.81$0.78$1.66$3.06$15.11
qwen3-omni-flash-2025-09-15non-thinking and thinking modes$0.43$3.81$0.78$1.66$3.06$15.11
qwen-omni-turbo Currently equivalent to qwen-omni-turbo-2025-03-26non-thinking mode$0.07$4.44$0.21$0.27$0.63$8.89
qwen-omni-turbo-latestnon-thinking mode$0.07$4.44$0.21$0.27$0.63$8.89
qwen-omni-turbo-2025-03-26non-thinking mode$0.07$4.44$0.21$0.27$0.63$8.89

Chinese mainland

If you select the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Input priceOutput price
Text/image/videoAudioText ** For multimodal inputText+audio** ** Audio billed only
qwen3.5-omni-plus Currently equivalent to qwen3.5-omni-plus-2026-03-15$0.96$7.29$5.5$29.29
qwen3.5-omni-plus-2026-03-15$0.96$7.29$5.5$29.29
qwen3.5-omni-flash Currently equivalent to qwen3.5-omni-flash-2026-03-15$0.3$2.48$1.83$9.9
qwen3.5-omni-flash-2026-03-15$0.3$2.48$1.83$9.9
More models
Model**ModeInput priceOutput price
TextAudioImage/videoText ** For text-only inputText** ** For multimodal inputText+audio** ** Audio billed only
qwen3-omni-flash Currently equivalent to qwen3-omni-flash-2025-12-01non-thinking and thinking modes$0.258$2.265$0.473$0.989$1.821$8.974
qwen3-omni-flash-2025-12-01non-thinking and thinking modes$0.258$2.265$0.473$0.989$1.821$8.974
qwen3-omni-flash-2025-09-15non-thinking and thinking modes$0.258$2.265$0.473$0.989$1.821$8.974
qwen-omni-turbo Currently equivalent to qwen-omni-turbo-2025-03-26non-thinking mode$0.058$3.584$0.216$0.230$0.646$7.168
qwen-omni-turbo-latestnon-thinking mode$0.058$3.584$0.216$0.230$0.646$7.168
qwen-omni-turbo-2025-03-26non-thinking mode$0.058$3.584$0.216$0.230$0.646$7.168
qwen-omni-turbo-2025-01-19non-thinking mode$0.058$3.584$0.216$0.230$0.646$7.168

Qwen-Omni-Realtime

Billing is based on input and output tokens. See billing and rate limits for how tokens are calculated for different modalities.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in your selected region. The supported region for this deployment scope is Asia Pacific SE 1 (Singapore).

Model name**Input price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
Text and imageAudioText ** Multimodal inputText + audio** ** Billed for audio only
qwen3.5-omni-plus-realtime$2.1$16.5$12.4$621 million input tokens and 1 million output tokensValid for 90 days after activating Model Studio.
qwen3.5-omni-plus-realtime-2026-03-15$2.1$16.5$12.4$62
qwen3.5-omni-flash-realtime$0.55$4.5$3.3$17.7
qwen3.5-omni-flash-realtime-2026-03-15$0.55$4.5$3.3$17.7

More models

Model name**Input price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
TextAudioImageText ** For text-only inputText** ** Multimodal inputText + audio** ** Billed for audio only
qwen3-omni-flash-realtime$0.52$4.57$0.94$1.99$3.67$18.131 million input tokens and 1 million output tokens (modality-agnostic)Valid for 90 days after activating Model Studio.
qwen3-omni-flash-realtime-2025-12-01$0.52$4.57$0.94$1.99$3.67$18.13
qwen3-omni-flash-realtime-2025-09-15$0.52$4.57$0.94$1.99$3.67$18.13
qwen-omni-turbo-realtime$0.270$4.440$0.840$1.070$2.520$8.890
qwen-omni-turbo-realtime-latest$0.270$4.440$0.840$1.070$2.520$8.890
qwen-omni-turbo-realtime-2025-05-08$0.270$4.440$0.840$1.070$2.520$8.890

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model name**Input price (per 1M tokens)Output price (per 1M tokens)
Text and imageAudioText ** Multimodal inputText + audio** ** Billed for audio only
qwen3.5-omni-plus-realtime$1.38$11$8.25$41.26
qwen3.5-omni-plus-realtime-2026-03-15$1.38$11$8.25$41.26
qwen3.5-omni-flash-realtime$0.45$3.71$2.75$14.71
qwen3.5-omni-flash-realtime-2026-03-15$0.45$3.71$2.75$14.71

More models

Model name**Input price (per 1M tokens)Output price (per 1M tokens)
TextAudioImageText ** For text-only inputText** ** Multimodal inputText + audio** ** Billed for audio only
qwen3-omni-flash-realtime$0.315$2.709$0.559$1.19$2.179$10.766
qwen3-omni-flash-realtime-2025-12-01$0.315$2.709$0.559$1.19$2.179$10.766
qwen3-omni-flash-realtime-2025-09-15$0.315$2.709$0.559$1.19$2.179$10.766
qwen-omni-turbo-realtime$0.230$3.584$0.861$0.918$2.581$7.168
qwen-omni-turbo-realtime-latest$0.230$3.584$0.861$0.918$2.581$7.168
qwen-omni-turbo-realtime-2025-05-08$0.230$3.584$0.861$0.918$2.581$7.168

QVQ

Billing is based on the number of input and output tokens. See Billing and rate limits for details on how tokens are calculated for different modalities.

International

If you select the International deployment scope, model inference resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

Model**Input priceOutput priceFree quota(Note)
qvq-max ** Currently equivalent to qvq-max-2025-03-25$1.2$4.81 million input and 1 million output tokens Valid for 90 days after activating Model Studio
qvq-max-latest$1.2$4.8
qvq-max-2025-03-25$1.2$4.8

Chinese mainland

If you select the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Input priceOutput price
qvq-max ** Currently equivalent to qvq-max-2025-03-25$1.147$4.588
qvq-max-latest$1.147$4.588
qvq-max-2025-05-15$1.147$4.588
qvq-max-2025-03-25$1.147$4.588
qvq-plus Currently equivalent to qvq-plus-2025-05-15$0.287$0.717
qvq-plus-latest$0.287$0.717
qvq-plus-2025-05-15$0.287$0.717

Qwen-VL

You are charged for input tokens and output tokens.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in Singapore, the only supported region for this scope.

Model name**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** Chain of Thought (CoT) + response**Free quota (Note)
qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.2$1.61 million tokens each for input and output Valid for 90 days after you activate Model Studio.
32K < tokens ≤ 128K$0.3$2.4
128K < tokens ≤ 256K$0.6$4.8
qwen3-vl-plus-2025-12-19thinking and non-thinking modes0 < tokens ≤ 32K$0.2$1.6
32K < tokens ≤ 128K$0.3$2.4
128K < tokens ≤ 256K$0.6$4.8
qwen3-vl-plus-2025-09-23thinking and non-thinking modes0 < tokens ≤ 32K$0.2$1.6
32K < tokens ≤ 128K$0.3$2.4
128K < tokens ≤ 256K$0.6$4.8
qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discount for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2026-01-22thinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2025-10-15thinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
More models
Model name**Input tokensInput price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
qwen-vl-max ** Currently equivalent to qwen-vl-max-2025-08-13 Discount for context cachingNo tiered pricing$0.8$3.21 million tokens each for input and outputValid for 90 days after you activate Model Studio.
qwen-vl-max-latestNo tiered pricing$0.8$3.2
qwen-vl-max-2025-08-13No tiered pricing$0.8$3.2
qwen-vl-max-2025-04-08No tiered pricing$0.8$3.2
qwen-vl-plus Currently equivalent to qwen-vl-plus-2025-08-15 Discount for context cachingNo tiered pricing$0.21$0.63
qwen-vl-plus-latestNo tiered pricing$0.21$0.63
qwen-vl-plus-2025-08-15No tiered pricing$0.21$0.63
qwen-vl-plus-2025-05-07No tiered pricing$0.21$0.63
qwen-vl-plus-2025-01-25No tiered pricing$0.21$0.63

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model**ModeInput tokensInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.143$1.434
32K < tokens ≤ 128K$0.215$2.15
128K < tokens ≤ 256K$0.43$4.301
qwen3-vl-plus-2025-09-23thinking and non-thinking modes0 < tokens ≤ 32K$0.143$1.434
32K < tokens ≤ 128K$0.215$2.15
128K < tokens ≤ 256K$0.43$4.301
qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2025-10-15 Discount for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.022$0.215
32K < tokens ≤ 128K$0.043$0.43
128K < tokens ≤ 256K$0.086$0.859
qwen3-vl-flash-2025-10-15thinking and non-thinking modes0 < tokens ≤ 32K$0.022$0.215
32K < tokens ≤ 128K$0.043$0.43
128K < tokens ≤ 256K$0.086$0.859

US

When the US deployment scope is selected, model inference resources are restricted to the US. Static data is stored in your selected region. The supported region is US (Virginia). Note

Models in the US deployment scope do not offer a free quota.

Model name**ModeInput tokens per requestInput price (per 1M tokens)Output price (per 1M tokens) ** Chain of thought + Response**
qwen3-vl-flash-us ** Discounts available for context cachingThinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2026-01-22-usThinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2025-10-15-usThinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96

Chinese mainland

When you select the Chinese mainland as the deployment scope, model inference compute resources are restricted to the Chinese mainland, and static data is stored in your selected region. China (Beijing) is the only supported region for this deployment scope. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**ModeInput tokensInput priceOutput price ** CoT and response**
qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discounts are available for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.143$1.434
32K < tokens ≤ 128K$0.215$2.15
128K < tokens ≤ 256K$0.43$4.301
qwen3-vl-plus-2025-12-19thinking and non-thinking modes0 < tokens ≤ 32K$0.143$1.434
32K < tokens ≤ 128K$0.215$2.15
128K < tokens ≤ 256K$0.43$4.301
qwen3-vl-plus-2025-09-23thinking and non-thinking modes0 < tokens ≤ 32K$0.143$1.434
32K < tokens ≤ 128K$0.215$2.15
128K < tokens ≤ 256K$0.43$4.301
qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discounts are available for context cachingthinking and non-thinking modes0 < tokens ≤ 32K$0.022$0.215
32K < tokens ≤ 128K$0.043$0.43
128K < tokens ≤ 256K$0.086$0.859
qwen3-vl-flash-2026-01-22thinking and non-thinking modes0 < tokens ≤ 32K$0.022$0.215
32K < tokens ≤ 128K$0.043$0.43
128K < tokens ≤ 256K$0.086$0.859
qwen3-vl-flash-2025-10-15thinking and non-thinking modes0 < tokens ≤ 32K$0.022$0.215
32K < tokens ≤ 128K$0.043$0.43
128K < tokens ≤ 256K$0.086$0.859
More models
Model**Input tokensInput priceOutput price
qwen-vl-max ** Currently equivalent to qwen-vl-max-2025-08-13 Discounts are available for context cachingNo tiered pricing$0.23$0.574
qwen-vl-max-latestNo tiered pricing$0.23$0.574
qwen-vl-max-2025-08-13No tiered pricing$0.23$0.574
qwen-vl-max-2025-04-08No tiered pricing$0.431$1.291
qwen-vl-max-2025-04-02No tiered pricing$0.431$1.291
qwen-vl-max-2025-01-25No tiered pricing$0.431$1.291
qwen-vl-max-2024-12-30No tiered pricing$0.431$1.291
qwen-vl-max-2024-11-19No tiered pricing$0.431$1.291
qwen-vl-plus Currently equivalent to qwen-vl-plus-2025-08-15 Discounts are available for context cachingNo tiered pricing$0.115$0.287
qwen-vl-plus-latestNo tiered pricing$0.115$0.287
qwen-vl-plus-2025-08-15No tiered pricing$0.115$0.287
qwen-vl-plus-2025-07-10No tiered pricing$0.022$0.216
qwen-vl-plus-2025-05-07No tiered pricing$0.216$0.646
qwen-vl-plus-2025-01-25No tiered pricing$0.216$0.646
qwen-vl-plus-2025-01-02No tiered pricing$0.216$0.646

China (Hong Kong)

Selecting the China (Hong Kong) deployment scope places both model inference compute resources and static data in the China (Hong Kong) region. Note

Models in the China (Hong Kong) deployment scope do not offer a free quota.

Model**ModeInput tokensInput priceOutput price ** CoT + answer**
qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context cachingnon-thought and thought modes0Model**ModeInput tokens
---------------
qwen3-vl-plus ** Discounts available for context cachingThinking and non-thinking modes0 < tokens ≤ 32K$0.2$1.6
32K < tokens ≤ 128K$0.3$2.4
128K < tokens ≤ 256K$0.6$4.8
qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discounts available for context cachingThinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2026-01-22Thinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96
qwen3-vl-flash-2025-10-15Thinking and non-thinking modes0 < tokens ≤ 32K$0.05$0.4
32K < tokens ≤ 128K$0.075$0.6
128K < tokens ≤ 256K$0.12$0.96

Qwen-OCR

You are charged for input tokens and output tokens.

International

Selecting the International deployment scope dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

Model**Input priceOutput priceFree quota (see note)
qwen-vl-ocr$0.07$0.161 million tokens for each model This quota expires 90 days after you activate Model Studio.
qwen-vl-ocr-2025-11-20

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

ModelInput priceOutput price
qwen-vl-ocr$0.043$0.072
qwen-vl-ocr-2025-11-20$0.043$0.072

Chinese mainland

Selecting the Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput priceOutput price
qwen-vl-ocr$0.043$0.072
qwen-vl-ocr-latest$0.043$0.072
qwen-vl-ocr-2025-11-20$0.043$0.072
qwen-vl-ocr-2025-08-28$0.717$0.717
qwen-vl-ocr-2025-04-13$0.717$0.717
qwen-vl-ocr-2024-10-28$0.717$0.717

Qwen-Math

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

Model nameInput priceOutput priceFree quota**(note)**
qwen-math-plus$0.574$1.721No free quota
qwen-math-plus-latest$0.574$1.721
qwen-math-plus-2024-09-19$0.574$1.721
qwen-math-plus-2024-08-16$0.574$1.721
qwen-math-turbo$0.287$0.861
qwen-math-turbo-latest$0.287$0.861
qwen-math-turbo-2024-09-19$0.287$0.861

Qwen-Coder

You are charged for input tokens and output tokens.

If the model supports context cache, only input tokens receive a discount.

International

In the International deployment scope, model inference runs on compute resources that are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in your selected region. The supported region for this scope is Singapore.

ModelInput tokensInput price (per 1M tokens)Output price (per 1M tokens)Free quota(note)
qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount0Model**Input tokensInput price (per 1M tokens)Output price (per 1M tokens)
------------
qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount0Model**Input tokensInput price (per 1M tokens)Output price (per 1M tokens)
------------
qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount0Model**Input tokensInput price (per 1M tokens)Output price (per 1M tokens)
------------
qwen-coder-plus ** Currently equivalent to qwen-coder-plus-2024-11-06no tiered pricing$0.502$1.004
qwen-coder-plus-latestno tiered pricing$0.502$1.004
qwen-coder-plus-2024-11-06no tiered pricing$0.502$1.004
qwen-coder-turbo Currently equivalent to qwen-coder-turbo-2024-09-19no tiered pricing$0.287$0.861
qwen-coder-turbo-latestno tiered pricing$0.287$0.861
qwen-coder-turbo-2024-09-19no tiered pricing$0.287$0.861

Qwen-MT

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the Singapore region.

Model**Input price (per 1M tokens)Output price (per 1M tokens)Free quota(note)
qwen-mt-plus$2.46$7.371M input and 1M output tokens Valid for 90 days after you activate Model Studio
qwen-mt-flash$0.16$0.49
qwen-mt-lite$0.12$0.36
qwen-mt-turbo$0.16$0.49

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

ModelInput price (per 1M tokens)Output price (per 1M tokens)
qwen-mt-plus$0.259$0.775
qwen-mt-flash$0.101$0.280
qwen-mt-lite$0.086$0.229

US

If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data is stored in your selected region. This deployment scope is available in the US (Virginia) region. Note

Models in the US deployment scope do not offer a free quota.

ModelInput price (per 1M tokens)Output price (per 1M tokens)
qwen-mt-lite-us$0.12$0.36

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput price (per 1M tokens)Output price (per 1M tokens)
qwen-mt-plus$0.259$0.775
qwen-mt-flash$0.101$0.280
qwen-mt-lite$0.086$0.229
qwen-mt-turbo$0.101$0.280

Qwen data mining

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

ModelInput priceOutput priceFree quota (Note)
qwen-doc-turbo$0.087$0.144No free quota

Qwen deep research

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

Model nameInput price (per million tokens)Output price (per million tokens)**Free quota **(Note)
qwen-deep-research$7.742$23.367None

Text generation - Qwen - Open source

Qwen3.6

You are charged for input tokens and output tokens.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

ModelInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode (chain of thought + response)
qwen3.6-35b-a3b0 < token ≤ 256K$0.248$1.485$1.485

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland), while static data is stored in your selected region. Supported region: Singapore (Singapore).

ModelInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)Free quota**(Note)**
Non-thinking modeThinking mode (chain of thought + response)
qwen3.6-35b-a3b0 < token ≤ 256K$0.375$2.25$2.251 million tokens for each model. Validity: 90 days after you activate Model Studio.
qwen3.6-27b0 < token ≤ 256K$0.6$3.6$3.6

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland, while static data is stored in your selected region. Supported region: China (Beijing).

ModelInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinking modeThinking mode (chain of thought + response)
qwen3.6-35b-a3b0 < token ≤ 256K$0.248$1.485$1.485
qwen3.6-27b0 < token ≤ 256K$0.412564$2.475384$2.475384

Qwen3.5

You are charged for input tokens and output tokens.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

Model nameInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinkingThinking (CoT + response)
qwen3.5-397b-a17b0 < token ≤ 128K$0.172$1.032$1.032
128K < token ≤ 256K$0.43$2.58$2.58
qwen3.5-122b-a10b0 < token ≤ 128K$0.115$0.917$0.917
128K < token ≤ 256K$0.287$2.294$2.294
qwen3.5-27b0 < token ≤ 128K$0.086$0.688$0.688
128K < token ≤ 256K$0.258$2.064$2.064
qwen3.5-35b-a3b0 < token ≤ 128K$0.057$0.459$0.459
128K < token ≤ 256K$0.229$1.835$1.835

International

If you select the International deployment scope, the platform dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data remains in your selected region. Supported region: Singapore.

Model nameInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)Free quota** (note)**
Non-thinkingThinking (CoT + response)
qwen3.5-397b-a17b0 < token ≤ 256K$0.6$3.6$3.61 million tokens per model This quota is valid for 90 days after you activate Model Studio.
qwen3.5-122b-a10b0 < token ≤ 256K$0.4$3.2$3.2
qwen3.5-27b0 < token ≤ 256K$0.3$2.4$2.4
qwen3.5-35b-a3b0 < token ≤ 256K$0.25$2$2

Chinese mainland

If you select the Chinese mainland deployment scope, the platform restricts model inference compute resources to the Chinese mainland. Static data remains in your selected region. Supported region: China (Beijing).

Model nameInput token rangeInput price (per 1M tokens)Output price (per 1M tokens)
Non-thinkingThinking (CoT + response)
qwen3.5-397b-a17b0 < token ≤ 128K$0.172$1.032$1.032
128K < token ≤ 256K$0.43$2.58$2.58
qwen3.5-122b-a10b0 < token ≤ 128K$0.115$0.917$0.917
128K < token ≤ 256K$0.287$2.294$2.294
qwen3.5-27b0 < token ≤ 128K$0.086$0.688$0.688
128K < token ≤ 256K$0.258$2.064$2.064
qwen3.5-35b-a3b0 < token ≤ 128K$0.057$0.459$0.459
128K < token ≤ 256K$0.229$1.835$1.835

Qwen3

You are charged for input tokens and output tokens.

International

With the International deployment scope, model inference resources are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in your selected region. This scope is available in the Singapore region.

ModelModeInput priceOutput priceFree quota (note)
Non-thinking modeThinking mode
qwen3-next-80b-a3b-thinkingThinking mode only$0.15-$1.21 million tokens per account. This quota is valid for 90 days after you activate Model Studio.
qwen3-next-80b-a3b-instructNon-thinking mode only$0.15$1.2-
qwen3-235b-a22b-thinking-2507Thinking mode only$0.23-$2.3
qwen3-235b-a22b-instruct-2507Non-thinking mode only$0.23$0.92-
qwen3-30b-a3b-thinking-2507Thinking mode only$0.2-$2.4
qwen3-30b-a3b-instruct-2507Non-thinking mode only$0.2$0.8-
qwen3-235b-a22bNon-thinking & thinking modes$0.7$2.8$8.4
qwen3-32bNon-thinking & thinking modes$0.16$0.64$0.64
qwen3-30b-a3bNon-thinking & thinking modes$0.2$0.8$2.4
qwen3-14bNon-thinking & thinking modes$0.35$1.4$4.2
qwen3-8bNon-thinking & thinking modes$0.18$0.7$2.1
qwen3-4bNon-thinking & thinking modes$0.11$0.42$1.26
qwen3-1.7bNon-thinking & thinking modes$0.11$0.42$1.26
qwen3-0.6bNon-thinking & thinking modes$0.11$0.42$1.26

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

ModelModeInput priceOutput price
Non-thinking modeThinking mode
qwen3-next-80b-a3b-thinkingThinking mode only$0.144-$1.434
qwen3-next-80b-a3b-instructNon-thinking mode only$0.144$0.574-
qwen3-235b-a22b-thinking-2507Thinking mode only$0.23-$2.3
qwen3-235b-a22b-instruct-2507Non-thinking mode only$0.23$0.92-
qwen3-30b-a3b-thinking-2507Thinking mode only$0.108-$1.076
qwen3-30b-a3b-instruct-2507Non-thinking mode only$0.108$0.431-
qwen3-235b-a22bNon-thinking & thinking modes$0.287$1.147$2.868
qwen3-32bNon-thinking & thinking modes$0.16$0.64$0.64
qwen3-30b-a3bNon-thinking & thinking modes$0.108$0.431$1.076
qwen3-14bNon-thinking & thinking modes$0.144$0.574$1.434
qwen3-8bNon-thinking & thinking modes$0.072$0.287$0.717

Chinese mainland

With the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland, and static data is stored in your selected region. This scope is available in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelModeInput priceOutput price
Non-thinking modeThinking mode
qwen3-next-80b-a3b-thinkingThinking mode only$0.144-$1.434
qwen3-next-80b-a3b-instructNon-thinking mode only$0.144$0.574-
qwen3-235b-a22b-thinking-2507Thinking mode only$0.287-$2.868
qwen3-235b-a22b-instruct-2507Non-thinking mode only$0.287$1.147-
qwen3-30b-a3b-thinking-2507Thinking mode only$0.108-$1.076
qwen3-30b-a3b-instruct-2507Non-thinking mode only$0.108$0.431-
qwen3-235b-a22bNon-thinking & thinking modes$0.287$1.147$2.868
qwen3-32bNon-thinking & thinking modes$0.287$1.147$2.868
qwen3-30b-a3bNon-thinking & thinking modes$0.108$0.431$1.076
qwen3-14bNon-thinking & thinking modes$0.144$0.574$1.434
qwen3-8bNon-thinking & thinking modes$0.072$0.287$0.717
qwen3-4bNon-thinking & thinking modes$0.044$0.173$0.431
qwen3-1.7bNon-thinking & thinking modes$0.044$0.173$0.431
qwen3-0.6bNon-thinking & thinking modes$0.044$0.173$0.431

QwQ - open source

You are charged for input tokens and output tokens.

ModelInput price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
qwq-32b$0.287$0.861No free quota

QwQ-Preview

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

ModelInput price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
qwq-32b-preview$0.287$0.861No free quota

Qwen2.5

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, compute resources for model inference are dynamically scheduled globally, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.

Model nameInput price (per million tokens)Output price (per million tokens)Free quota (note)
qwen2.5-14b-instruct-1m$0.805$3.221 million tokens each for input and output Valid for 90 days after you activate Model Studio
qwen2.5-7b-instruct-1m$0.368$1.47
qwen2.5-72b-instruct$1.4$5.6
qwen2.5-32b-instruct$0.7$2.8
qwen2.5-14b-instruct$0.35$1.4
qwen2.5-7b-instruct$0.175$0.7

Chinese mainland

If you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput price (per million tokens)Output price (per million tokens)
qwen2.5-14b-instruct-1m$0.144$0.431
qwen2.5-7b-instruct-1m$0.072$0.144
qwen2.5-72b-instruct$0.574$1.721
qwen2.5-32b-instruct$0.287$0.861
qwen2.5-14b-instruct$0.144$0.431
qwen2.5-7b-instruct$0.072$0.144
qwen2.5-3b-instruct$0.044$0.130
qwen2.5-1.5b-instructFree for a limited time
qwen2.5-0.5b-instruct

QVQ

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

ModelInput price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
qvq-72b-preview$1.721$5.161No free quota

Qwen-Omni

You are billed for input and output tokens. For details on how tokens are counted for different modalities, see billing and rate limiting.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the Singapore region.

ModelInput priceOutput priceFree quota (Note)
TextAudioImage/videoText ** Text-only inputText** ** Multimodal inputText+audio** ** Billed for audio only
qwen2.5-omni-7b$0.10$6.76$0.28$0.40$0.84$13.511 million tokens (modality-agnostic)Valid for 90 days after activating Model Studio.

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Input priceOutput price
TextAudioImage/videoText ** Text-only inputText** ** Multimodal inputText+audio** ** Billed for audio only
qwen2.5-omni-7b$0.087$5.448$0.287$0.345$0.861$10.895

Qwen3-Omni-Captioner

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.

Model name**Input price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
qwen3-omni-30b-a3b-captioner$3.81$3.061 million tokens Valid for 90 days after activating Model Studio.

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput price (per 1M tokens)Output price (per 1M tokens)
qwen3-omni-30b-a3b-captioner$2.265$1.821

Qwen-VL

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide, excluding the Chinese mainland. Static data resides in your selected region. The supported region is Singapore.

Model nameModeInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**Free quota (note)
qwen3-vl-235b-a22b-thinkingthinking mode$0.4$41 million tokens for each model. The quota is valid for 90 days after you activate Model Studio.
qwen3-vl-235b-a22b-instructnon-thinking mode$0.4$1.6
qwen3-vl-32b-thinkingthinking mode$0.16$0.64
qwen3-vl-32b-instructnon-thinking mode$0.16$0.64
qwen3-vl-30b-a3b-thinkingthinking mode$0.2$2.4
qwen3-vl-30b-a3b-instructnon-thinking mode$0.2$0.8
qwen3-vl-8b-thinkingthinking mode$0.18$2.1
qwen3-vl-8b-instructnon-thinking mode$0.18$0.7
More models
Model nameInput price (per 1M tokens)Output price (per 1M tokens)Free quota (note)
qwen2.5-vl-72b-instruct$2.8$8.41 million tokens for each model. The quota is valid for 90 days after you activate Model Studio.
qwen2.5-vl-32b-instruct$1.4$4.2
qwen2.5-vl-7b-instruct$0.35$1.05
qwen2.5-vl-3b-instruct$0.21$0.63

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameModeInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3-vl-235b-a22b-thinkingthinking mode$0.287$2.867
qwen3-vl-235b-a22b-instructnon-thinking mode$0.287$1.147
qwen3-vl-32b-thinkingthinking mode$0.16$0.64
qwen3-vl-32b-instructnon-thinking mode$0.16$0.64
qwen3-vl-30b-a3b-thinkingthinking mode$0.108$1.076
qwen3-vl-30b-a3b-instructnon-thinking mode$0.108$0.431
qwen3-vl-8b-thinkingthinking mode$0.072$0.717
qwen3-vl-8b-instructnon-thinking mode$0.072$0.287

Chinese mainland

If you select the Chinese mainland deployment scope, compute resources for model inference operate exclusively within the Chinese mainland. Static data resides in your selected region. The supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameModeInput price (per 1M tokens)Output price (per 1M tokens) ** CoT + response**
qwen3-vl-235b-a22b-thinkingthinking mode$0.287$2.867
qwen3-vl-235b-a22b-instructnon-thinking mode$0.287$1.147
qwen3-vl-32b-thinkingthinking mode$0.287$2.868
qwen3-vl-32b-instructnon-thinking mode$0.287$1.147
qwen3-vl-30b-a3b-thinkingthinking mode$0.108$1.076
qwen3-vl-30b-a3b-instructnon-thinking mode$0.108$0.431
qwen3-vl-8b-thinkingthinking mode$0.072$0.717
qwen3-vl-8b-instructnon-thinking mode$0.072$0.287
More models
Model nameInput price (per 1M tokens)Output price (per 1M tokens)
qwen2.5-vl-72b-instruct$2.294$6.881
qwen2.5-vl-32b-instruct$1.147$3.441
qwen2.5-vl-7b-instruct$0.287$0.717
qwen2.5-vl-3b-instruct$0.173$0.517
qwen2-vl-72b-instruct$2.294$6.881
qwen2-vl-7b-instructFree for a limited time
qwen2-vl-2b-instruct

Qwen-Math

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

ModelInput priceOutput priceFree quota (Note)
qwen2.5-math-72b-instruct$0.574$1.721No free quota
qwen2.5-math-7b-instruct$0.144$0.287
qwen2.5-math-1.5b-instructFree for a limited time

Qwen-Coder

You are charged for input tokens and output tokens.

International

For the International deployment scope, model inference resources are scheduled dynamically worldwide (excluding the Chinese mainland). Your static data is stored in the selected region. The only supported region for this scope is Singapore.

Model nameInput tokensInput price (per million tokens)Output price (per million tokens)Free quota (note)
qwen3-coder-next0 < tokens ≤ 32K$0.3$1.51 million input tokens and 1 million output tokens Validity: Within 90 days of activating Model Studio
32K < tokens ≤ 128K$0.5$2.5
128K < tokens ≤ 256K$0.8$4
qwen3-coder-480b-a35b-instruct0 < tokens ≤ 32K$1.5$7.5
32K < tokens ≤ 128K$2.7$13.5
128K < tokens ≤ 200K$4.5$22.5
qwen3-coder-30b-a3b-instruct0 < tokens ≤ 32K$0.45$2.25
32K < tokens ≤ 128K$0.75$3.75
128K < tokens ≤ 200K$1.2$6

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameInput tokensInput price (per million tokens)Output price (per million tokens)
qwen3-coder-480b-a35b-instruct0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.291$5.161
128K < tokens ≤ 200K$2.151$8.602
qwen3-coder-30b-a3b-instruct0 < tokens ≤ 32K$0.216$0.861
32K < tokens ≤ 128K$0.323$1.291
128K < tokens ≤ 200K$0.538$2.151

Chinese mainland

For the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Your static data is stored in the selected region. The only supported region for this scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput tokensInput price (per million tokens)Output price (per million tokens)
qwen3-coder-next0 < tokens ≤ 32K$0.144$0.574
32K < tokens ≤ 128K$0.216$0.861
128K < tokens ≤ 256K$0.359$1.434
qwen3-coder-480b-a35b-instruct0 < tokens ≤ 32K$0.861$3.441
32K < tokens ≤ 128K$1.291$5.161
128K < tokens ≤ 200K$2.151$8.602
qwen3-coder-30b-a3b-instruct0 < tokens ≤ 32K$0.216$0.861
32K < tokens ≤ 128K$0.323$1.291
128K < tokens ≤ 200K$0.538$2.151
qwen2.5-coder-32b-instructNo tiered pricing$0.287$0.861
qwen2.5-coder-14b-instructNo tiered pricing$0.287$0.861
qwen2.5-coder-7b-instructNo tiered pricing$0.144$0.287
qwen2.5-coder-3b-instructNo tiered pricingfree for a limited time
qwen2.5-coder-1.5b-instructNo tiered pricing
qwen2.5-coder-0.5b-instructNo tiered pricing

European Union

For the European Union deployment scope, model inference resources are restricted to the European Union. Your static data is stored in the selected region. For this scope, the only supported region is Germany (Frankfurt). Note

Models in the EU deployment scope do not offer a free quota.

Model nameInput tokensInput price (per million tokens)Output price (per million tokens)
qwen3-coder-next0 < tokens ≤ 32K$0.3$1.5
32K < tokens ≤ 128K$0.5$2.5
128K < tokens ≤ 256K$0.8$4

Text generation - third-party models

DeepSeek

You are charged for input tokens and output tokens.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).

Model nameInput price (per 1M tokens)Output price (per 1M tokens)Free quota
deepseek-v4-pro ** Discounts are available for context caching.$1.65$3.3No free quota
deepseek-v4-flash Discounts are available for context caching.$0.14$0.28

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland), while static data is stored in your selected region. The supported region for this deployment scope is Singapore.

Model name**Input price (per 1M tokens)Output price (per 1M tokens)Free quota
deepseek-v4-pro ** Discounts are available for context caching.$2.400$4.8001 million tokensValidity period: 90 days after activating Model Studio
deepseek-v4-flash Discounts are available for context caching.$0.200$0.4001 million tokensValidity period: 90 days after activating Model Studio
deepseek-v3.2 Discounts are available for context caching.$0.57$1.711 million tokensValidity period: 90 days after activating Model Studio

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland, while static data is stored in your selected region. The supported region for this deployment scope is China (Beijing).

Model name**Input price (per 1M tokens)Output price (per 1M tokens)Free quota (Note)
deepseek-v4-pro ** Discounts are available for context caching.$1.65$3.301No free quota
deepseek-v4-flash Discounts are available for context caching.$0.138$0.275
deepseek-v3.2 Discounts are available for context caching.$0.287$0.431
deepseek-v3.2-exp$0.287$0.431
deepseek-v3.1$0.574$1.721
deepseek-r1$0.574$2.294
deepseek-r1-0528$0.574$2.294
deepseek-v3$0.287$1.147
deepseek-r1-distill-qwen-1.5bFree for a limited time
deepseek-r1-distill-qwen-7b$0.072$0.144No free quota
deepseek-r1-distill-qwen-14b$0.144$0.431
deepseek-r1-distill-qwen-32b$0.287$0.861
deepseek-r1-distill-llama-8bFree for a limited time
deepseek-r1-distill-llama-70bFree for a limited time

Kimi

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

Model name**ModeInput price (per 1M tokens)Output price (per 1M tokens)Free quota(Note)
kimi-k2.6thinking and non-thinking modes$0.8939$3.7131No free quota
kimi-k2.5thinking and non-thinking modes$0.574$3.011
kimi-k2-thinkingthinking mode only$0.574$2.294
Moonshot-Kimi-K2-Instructnon-thinking mode$0.574$2.294

MiniMax

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.

Model nameModeInput price (per 1M tokens)Output price (per 1M tokens) ** Chain of Thought (CoT) and response**
MiniMax-M2.5thinking mode only$0.304$1.213

GLM

You are charged for input tokens and output tokens.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameModelRequest input tokensInput price (per 1M tokens)Output price (per 1M tokens) ** Chain of Thought (CoT) and response**
glm-5.1thinking and non-thinking modes0 < tokens ≤ 32K$0.825$3.301
32K < tokens ≤ 200K$1.100$3.851

Chinese mainland

With the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland, and static data is stored in your selected region. This scope is available in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameModelRequest input tokensInput price (per 1M tokens)Output price (per 1M tokens) ** Chain of Thought (CoT) and response**
glm-5.1thinking and non-thinking modes0 < tokens ≤ 32K$0.825$3.301
32K < tokens ≤ 200K$1.100$3.851
glm-5thinking and non-thinking modes0 < tokens ≤ 32K$0.573$2.58
32K < tokens ≤ 166K$0.86$3.154
glm-4.7thinking and non-thinking modes0 < tokens ≤ 32K$0.431$2.007
32K < tokens ≤ 166K$0.574$2.294
glm-4.6thinking and non-thinking modes0 < tokens ≤ 32K$0.431$2.007
32K < tokens ≤ 166K$0.574$2.294

Image generation

You are not charged for input. You are charged for output based on the number of successfully generated images.

Formula: Cost = Image unit price × Number of images generated.

Notes:

  • Cost does not depend on image resolution or aspect ratio.

  • Failed requests incur no cost and do not consume your free quota.

Billing example: Some images fail to generate Assume the image unit price is $0.10 per image. If you call the API to generate four images but only three image URLs return successfully, the system charges only for the three successfully generated images.

  • Number billed: 3 images.

  • Cost calculation: 0.1 × 3 = $0.3.

Qwen text-to-image

You are charged for output only. For billing rules, see Image generation.

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

ModelOutput priceFree quota (Note)
qwen-image-2.0-pro$0.075/image100 images per model. The free quota is valid for 90 days after you activate Model Studio.
qwen-image-2.0-pro-2026-04-22$0.075/image
qwen-image-2.0-pro-2026-03-03$0.075/image
qwen-image-2.0$0.035/image
qwen-image-2.0-2026-03-03$0.035/image
qwen-image-max ** Currently equivalent to qwen-image-max-2025-12-30$0.075/image
qwen-image-max-2025-12-30$0.075/image
qwen-image-plus Currently equivalent to qwen-image$0.03/image
qwen-image-plus-2026-01-09$0.03/image
qwen-image$0.035/image

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Output price
qwen-image-2.0-pro$0.071676/image
qwen-image-2.0-pro-2026-04-22$0.071676/image
qwen-image-2.0-pro-2026-03-03$0.071676/image
qwen-image-2.0$0.028671/image
qwen-image-2.0-2026-03-03$0.028671/image
qwen-image-max ** Currently equivalent to qwen-image-max-2025-12-30$0.071677/image
qwen-image-max-2025-12-30$0.071677/image
qwen-image-plus Currently equivalent to qwen-image$0.028671/image
qwen-image-plus-2026-01-09$0.028671/image
qwen-image$0.035/image

Qwen-Image-Edit

You are charged for output only. For billing rules, see Image generation.

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

Model**Output priceFree quota (Note)
qwen-image-2.0-pro$0.075/image100 images per model. The free quota is valid for 90 days after you activate Model Studio.
qwen-image-2.0-pro-2026-04-22$0.075/image
qwen-image-2.0-pro-2026-03-03$0.075/image
qwen-image-2.0$0.035/image
qwen-image-2.0-2026-03-03$0.035/image
qwen-image-edit-max ** Currently equivalent to qwen-image-edit-max-2026-01-16$0.075/image
qwen-image-edit-max-2026-01-16$0.075/image
qwen-image-edit-plus Currently equivalent to qwen-image-edit-plus-2025-10-30$0.03/image
qwen-image-edit-plus-2025-12-15$0.03/image
qwen-image-edit-plus-2025-10-30$0.03/image
qwen-image-edit$0.045/image

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model**Output price
qwen-image-2.0-pro$0.071676/image
qwen-image-2.0-pro-2026-04-22$0.071676/image
qwen-image-2.0-pro-2026-03-03$0.071676/image
qwen-image-2.0$0.028671/image
qwen-image-2.0-2026-03-03$0.028671/image
qwen-image-edit-max ** Currently equivalent to qwen-image-edit-max-2026-01-16$0.071677/image
qwen-image-edit-max-2026-01-16$0.071677/image
qwen-image-edit-plus Currently equivalent to qwen-image-edit-plus-2025-10-30$0.028671/image
qwen-image-edit-plus-2025-12-15$0.028671/image
qwen-image-edit-plus-2025-10-30$0.028671/image
qwen-image-edit$0.043/image

Qwen-MT-Image

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

You are charged for output only. For billing rules, see Image generation.

Model**Output priceFree quota (Note)
qwen-mt-image$0.000431/imageNo free quota

Qwen-Z-Image text-to-image

You are charged for output only. For billing rules, see Image generation.

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

ModelOutput priceFree quota (Note)
z-image-turboDisabling prompt rewrite (prompt_extend=false): $0.015/imageEnabling prompt rewrite (prompt_extend=true): $0.03/image100 imagesValidity period: 90 days after Alibaba Cloud Model Studio is activated.

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput price
z-image-turboPrompt rewrite disabled (prompt_extend=false): $0.01434 per imageEnabling prompt rewrite (prompt_extend=true): $0.02868 per image

Wan text-to-image

You are charged for output only. For billing rules, see Image generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

ModelOutput price
wan2.6-t2i$0.028671/image

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

ModelOutput priceFree quota (Note)
wan2.6-t2i$0.03/image50 images
wan2.5-t2i-preview$0.03/image50 images
wan2.2-t2i-plus$0.05/image100 images
wan2.2-t2i-flash$0.025/image100 images
wan2.1-t2i-plus$0.05/image200 images
wan2.1-t2i-turbo$0.025/image200 images

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput price
wan2.6-t2i$0.028671/image
wan2.5-t2i-preview$0.028671/image
wan2.2-t2i-plus$0.020070/image
wan2.2-t2i-flash$0.028671/image
wanx2.1-t2i-plus$0.028671/image
wanx2.1-t2i-turbo$0.020070/image
wanx2.0-t2i-turbo$0.005735/image

Wan image generation and editing

You are charged for output only. For billing rules, see Image generation.

Global (Virginia)

Note

Models deployed globally (Virginia) do not offer a free quota.

ModelOutput price
wan2.6-image$0.028671/image

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

ModelOutput priceFree quota (Note)
wan2.7-image-pro$0.075/image50 images. The free quota is valid for 90 days after you activate Model Studio.
wan2.7-image$0.03/image50 images. The free quota is valid for 90 days after you activate Model Studio.
wan2.6-image$0.03/image50 images. The free quota is valid for 90 days after you activate Model Studio.

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput price
wan2.7-image-pro$0.068761/image
wan2.7-image$0.028671/image
wan2.6-image$0.028671/image

Wan general image editing

You are charged for output only. For billing rules, see Image generation.

International

When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.

Model serviceModelOutput priceFree quota (Note)
General Image Editing 2.5wan2.5-i2i-preview$0.03/image50 images. The free quota is valid for 90 days after you activate Model Studio.

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model serviceModelOutput price
General Image Editing 2.5wan2.5-i2i-preview$0.028671/image
General Image Editing 2.1wanx2.1-imageedit$0.020070/image

OutfitAnyone

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • aitryon-plus: You are charged for output only. For billing rules, see Image generation.

  • aitryon-parsing-v1: You are charged per input image, not per output image. Failed requests do not incur charges.

Model serviceModelPriceFree quota (Note)
OutfitAnyone - Plusaitryon-plus$0.071677/imageNo free quota
OutfitAnyone - Image Parsingaitryon-parsing-v1$0.000574/image

Video generation

You are not charged for input. You are charged for output based on the total duration of successfully generated videos (in seconds).

Formula: Cost = Video unit price × Video duration (seconds).

Notes:

  • Some models charge by output video resolution. Prices differ for resolutions such as 480P, 720P, and 1080P.

  • Some models charge by output video edition. Prices differ for editions such as Standard Edition and Professional Edition.

  • Some models charge by output video aspect ratio. Prices differ for aspect ratios such as 1:1 and 3:4.

  • Some models use a flat rate, regardless of resolution, edition, or aspect ratio.

  • Failed requests incur no cost and do not consume your free quota.

HappyHorse - Text-to-video

Billing is based on output only. For billing rules, see video generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

ModelOutput resolutionUnit price
happyhorse-1.0-t2v720p$0.123769/second
1080p$0.220034/second

International

In the International deployment scope, model inference compute resources are dynamically scheduled globally, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.

ModelOutput resolutionUnit priceFree Tier (Note) Valid for 90 days after activating Model Studio
happyhorse-1.0-t2v720p$0.14/second10 seconds
1080p$0.24/second

Chinese mainland

In the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput resolutionUnit price
happyhorse-1.0-t2v720p$0.123769/second
1080p$0.220034/second

HappyHorse - image-to-video (first frame)

You are charged for output only. For billing rules, see video generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
happyhorse-1.0-i2v720P$0.123769/second
1080P$0.220034/second

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in Singapore, the only supported region for this scope.

Model nameOutput video resolutionUnit priceFree quota (Note) Valid for 90 days after you activate Model Studio.
happyhorse-1.0-i2v720P$0.14/second10 seconds
1080P$0.24/second

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in China (Beijing), the only supported region for this scope. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
happyhorse-1.0-i2v720P$0.123769/second
1080P$0.220034/second

HappyHorse - Reference to video

You are billed only for the output. For billing rules, see video generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
happyhorse-1.0-r2v720p$0.123769/second
1080p$0.220034/second

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the region you select. Supported region: Singapore.

Model nameOutput video resolutionUnit priceFree quota(Note) Validity: 90 days after activating Model Studio
happyhorse-1.0-r2v720p$0.14/second10 seconds
1080p$0.24/second

Chinese mainland

When the deployment scope is Mainland China, model inference computing resources are limited to Mainland China, and static data is stored in the region that you select. The supported region for this deployment scope is: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
happyhorse-1.0-r2v720p$0.123769/second
1080p$0.220034/second

HappyHorse - video editing

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Billing rules: You are billed for both input and output videos per second. Failed requests are not billed.

ModelOutput video resolutionUnit price
happyhorse-1.0-video-edit720p$0.123769/second
1080p$0.220034/second

International

The International deployment scope dynamically schedules model inference compute resources worldwide (excluding the Chinese mainland). Static data is stored in the only supported region: Singapore.

Billing rules: You are billed for both input and output videos per second. Failed requests are not billed and do not count against your free quota.

ModelOutput video resolutionUnit priceFree quota (Note) Validity: 90 days after you activate Model Studio
happyhorse-1.0-video-edit720p$0.14/second10 seconds
1080p$0.24/second

Chinese mainland

The Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in the only supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Billing rules: You are billed for both input and output videos per second. Failed requests are not billed.

ModelOutput video resolutionUnit price
happyhorse-1.0-video-edit720p$0.123769/second
1080p$0.220034/second

Wan - text-to-video

You are billed for output only. For billing rules, see video generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
wan2.6-t2v720p$0.086012/second
1080p$0.143353/second

International

When you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in the Singapore (Singapore) region.

Model nameOutput video resolutionUnit priceFree quota (Note) Validity: 90 days after you activate Model Studio
wan2.7-t2v-2026-04-25720p$0.10/second50 seconds
1080p$0.15/second
wan2.7-t2v720p$0.10/second50 seconds
1080p$0.15/second
wan2.6-t2v720p$0.10/second50 seconds
1080p$0.15/second
wan2.5-t2v-preview480p$0.05/second50 seconds
720p$0.10/second
1080p$0.15/second
wan2.2-t2v-plus480p$0.02/second50 seconds
1080p$0.10/second
wan2.1-t2v-turbo480p$0.036/second200 seconds
720p$0.036/second
wan2.1-t2v-plus720p$0.10/second200 seconds

US

When you select the US deployment scope, compute resources for model inference are restricted to the United States. Static data is stored in the US (Virginia) region. Note

Models in the US deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
wan2.6-t2v-us720p$0.1/second
1080p$0.15/second

Chinese mainland

When you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland. Static data is stored in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput video resolutionUnit price
wan2.7-t2v-2026-04-25720p$0.086012/second
1080p$0.143353/second
wan2.7-t2v720p$0.086012/second
1080p$0.143353/second
wan2.6-t2v720p$0.086012/second
1080p$0.143353/second
wan2.5-t2v-preview480p$0.043006/second
720p$0.086012/second
1080p$0.143353/second
wan2.2-t2v-plus480p$0.02007/second
1080p$0.100347/second
wan2.1-t2v-turbo480p$0.034405/second
720p$0.034405/second
wan2.1-t2v-plus720p$0.100347/second

Wan - image-to-video

You are billed for output only. For billing rules, see video generation.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.

ModelOutput video typeOutput video resolutionPriceFree quota (Note) Valid for 90 days after activating Model Studio
wan2.7-i2v-2026-04-25Video with audio720p$0.10/second50 seconds
1080p$0.15/second
wan2.7-i2vVideo with audio720p$0.10/second50 seconds
1080p$0.15/second

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput video typeOutput video resolutionPrice
wan2.7-i2v-2026-04-25Video with audio720p$0.086012/second
1080p$0.143353/second
wan2.7-i2vVideo with audio720p$0.086012/second
1080p$0.143353/second

Wan: Image-to-video (from first frame)

You are billed for output only. For billing rules, see video generation.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameOutput typeOutput video resolutionUnit price
wan2.6-i2vvideo with audio720p$0.086012/second
1080p$0.143353/second

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data resides in your selected region. Supported region: Singapore.

Model nameOutput typeOutput video resolutionUnit priceFree quota (Note) Valid for 90 days after activating Model Studio
wan2.6-i2v-flashvideo with audioaudio=true720p$0.05/second50 seconds
1080p$0.075/second
video without audioaudio=false720p$0.025/second
1080p$0.0375/second
wan2.6-i2vvideo with audio720p$0.10/second50 seconds
1080p$0.15/second
wan2.5-i2v-previewvideo with audio480p$0.05/second50 seconds
720p$0.10/second
1080p$0.15/second
wan2.2-i2v-flashvideo without audio480p$0.015/second50 seconds
720p$0.036/second
wan2.2-i2v-plusvideo without audio480p$0.02/second50 seconds
1080p$0.10/second
wan2.1-t2v-turbovideo without audio480p$0.036/second200 seconds
720p$0.036/second
wan2.1-t2v-plusvideo without audio720p$0.10/second200 seconds

US

If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data resides in your selected region. Supported region: US (Virginia). Note

Models in the US deployment scope do not offer a free quota.

Model nameOutput typeOutput video resolutionUnit price
wan2.6-i2v-usvideo with audio720p$0.1/second
1080p$0.15/second

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data resides in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput typeOutput video resolutionUnit price
wan2.6-i2v-flashvideo with audioaudio=true720p$0.043006/second
1080p$0.071676/second
video without audioaudio=false720p$0.021503/second
1080p$0.035838/second
wan2.6-i2vvideo with audio720p$0.086012/second
1080p$0.143353/second
wan2.5-i2v-previewvideo with audio480p$0.043006/second
720p$0.086012/second
1080p$0.143353/second
wan2.2-i2v-plusvideo without audio480p$0.02007/second
1080p$0.100347/second
wanx2.1-t2v-turbovideo without audio480p$0.034405/second
720p$0.034405/second
wanx2.1-t2v-plusvideo without audio720p$0.100347/second

Wan - Image-to-Video (Start and End Frames)

You are billed for output only. For billing rules, see video generation.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data resides in the region you select. The supported region for this deployment scope is Singapore.

ModelOutput resolutionOutput priceFree quota (Note) Valid for 90 days after you activate Alibaba Cloud Model Studio
wan2.2-kf2v-flash480p$0.015/second50 seconds
720p$0.036/second
1080p$0.07/second
wan2.1-kf2v-plus720p$0.10/second200 seconds

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data resides in the region you select. The supported region for this deployment scope is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput resolutionOutput price
wan2.2-kf2v-flash480p$0.014335/second
720p$0.028671/second
1080p$0.068809/second
wanx2.1-kf2v-plus720p$0.100347/second

Wan - reference-to-video

Billing rules: You are billed for both input and output videos based on their duration in seconds. Failed requests do not incur charges or consume your free quota.

Billing formula: Billed duration = Input video duration (up to 5 seconds) + Output video duration.

  • The billed duration for the input video is capped at 5 seconds . For calculation rules, see Billing and rate limits.

  • The billed duration for the output video is the duration in seconds of the successfully generated video.

Global

If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note

Models in the global deployment scope do not offer a free quota.

Model nameOutput video typeOutput video resolutionInput and output price
wan2.6-r2vvideo with audio720p$0.086012/second
1080p$0.143353/second

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

Model nameOutput video typeOutput video resolutionInput and output priceFree quota (Note) Validity: 90 days after activating Model Studio
wan2.7-r2vvideo with audio720p$0.10/second50 seconds
1080p$0.15/second
wan2.6-r2v-flashvideo with audioaudio=true720p$0.05/second50 seconds
1080p$0.075/second
video without audioaudio=false720p$0.025/second
1080p$0.0375/second
wan2.6-r2vvideo with audio720p$0.10/second50 seconds
1080p$0.15/second

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput video typeOutput video resolutionInput and output price
wan2.7-r2vvideo with audio720p$0.086012/second
1080p$0.143353/second
wan2.6-r2v-flashvideo with audioaudio=true720p$0.043006/second
1080p$0.071676/second
video without audioaudio=false720p$0.021503/second
1080p$0.035838/second
wan2.6-r2vvideo with audio720p$0.086012/second
1080p$0.143353/second

Wan - video editing

International

When the deployment scope is International, model inference computing resources are dynamically scheduled globally (excluding mainland China). Static data is stored in the region you select. The supported region for this deployment scope is Singapore.

You are charged for both input and output videos per second of video duration. Failed tasks are not billed and do not consume your free quota.

Model nameOutput video resolutionCombined priceFree quota (Note) Valid for 90 days after you activate Model Studio.
wan2.7-videoedit720p$0.10/second50 seconds
1080p$0.15/second

You are charged for output videos per second of video duration. Failed tasks are not billed and do not consume your free quota.

Model nameOutput video resolutionOutput priceFree quota (Note) Valid for 90 days after you activate Model Studio.
wan2.1-vace-plus720p$0.10/second50 seconds

Chinese mainland

When the deployment scope is mainland China, model inference computing resources are restricted to mainland China, and static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

You are charged for both input and output videos per second of video duration. You are not charged for failed tasks.

Model nameOutput video resolutionCombined price
wan2.7-videoedit720p$0.086012/second
1080p$0.143353/second

You are charged for output videos per second of video duration. You are not charged for failed tasks.

Model nameOutput video resolutionOutput price
wanx2.1-vace-plus720p$0.100347/second

Wan - digital human

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • wan2.2-s2v-detect: You are charged for each input image per successful request, regardless of the detection result. There is no charge for output.

  • wan2.2-s2v: You are billed per second for each successfully generated output video. There is no charge for input. For billing rules, see video generation.

ServiceModel namePriceFree quota (Note)
Image detectionwan2.2-s2v-detectInput image: $0.000574/imageNo free quota
Video generationwan2.2-s2vOutput video: - 480P: $0.071677/second - 720P: $0.129018/second

Wan - image-to-motion

You are billed for output only. For billing rules, see video generation.

International

If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Singapore is the only supported region for this deployment scope.

ModelOutput video modeOutput priceFree quotaRules
wan2.2-animate-movestandard mode wan-std$0.12/second50 secondsValid for 90 days after you activate Model Studio.
professional mode wan-pro$0.18/second

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. China (Beijing) is the only supported region for this deployment scope. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelOutput video modeOutput price
wan2.2-animate-movestandard mode wan-std$0.06/second
professional mode wan-pro$0.09/second

Wan - video character swap

You are billed for output only. For billing rules, see video generation.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in the selected region, which must be Singapore.

Model nameOutput modePriceFree Tier (Note)
wan2.2-animate-mixStandard mode wan-std$0.18/second50 secondsThe free quota is valid for 90 days after you activate Model Studio.
Professional modewan-pro$0.26/second

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in the selected region, which must be China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameOutput modePrice
wan2.2-animate-mixStandard mode wan-std$0.09/second
Professional modewan-pro$0.13/second

AnimateAnyone

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • animate-anyone-detect-gen2: Charges are based on the input image. You are billed once for each input image per successful request, regardless of the detection outcome.

  • animate-anyone-template-gen2: Charges are based on the output video. You are billed per second for each successfully generated video. For billing rules, see video generation.

  • animate-anyone-gen2: Charges are based on the output video. You are billed per second for each successfully generated video. For billing rules, see video generation.

Model serviceModelPriceFree quota(Note)
Image detectionanimate-anyone-detect-gen2$0.000574 per input imageNo free quota
Action template generationanimate-anyone-template-gen2Output video: $0.011469/second
Video generationanimate-anyone-gen2Output video: $0.011469/second

EMO

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • emo-detect-v1: Billing is based on input, not output. You are charged for each input image per successful request, regardless of the detection result.

  • emo-v1: Billing is based on output, not input. You are charged per second of successfully generated video. For billing rules, see video generation.

Model serviceModel namePriceFree quota(note)
image detectionemo-detect-v1input image: $0.000574/imageNo free quota
video generationemo-v1output video: - Video (1:1 aspect ratio): $0.011469/second - Video (3:4 aspect ratio): $0.022937/second

LivePortrait

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • liveportrait-detect: You are charged for input but not for output. Each successful request is charged once per input image, regardless of the detection result.

  • liveportrait: You are charged for output but not for input. You are billed per second for each successfully generated video. For more information about billing rules, see video generation.

ServiceModelPriceFree quota (Note)
Image detectionliveportrait-detectInput image: $0.000574 per imageNo free quota is provided.
Video generationliveportraitOutput video: $0.002868 per second

Emoji

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

  • emoji-detect-v1: You are charged for the input, not the output. This charge applies once per input image for each successful request, regardless of the detection result.

  • emoji-v1: You are not charged for the input. You are billed per second of successfully generated video. For billing rules, see video generation.

Model serviceModel nameUnit priceFree quota (Note)
Image detectionemoji-detect-v1$0.000574 per input imageNo free quota
Video generationemoji-v1$0.011469 per second

VideoRetalk

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

You are billed for output only. For billing rules, see video generation.

Model nameOutput priceFree quota (Note)
videoretalk$0.011469/secondNo free quota

Video style transform

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

You are billed for output only. For billing rules, see video generation.

Model nameOutput resolutionUnit priceFree quota Note
video-style-transform540p$0.028671/secondNo free quota
720p$0.071677/second

Music generation

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Billing rule: You are billed only for the duration (in seconds) of the output audio.

ModelUnit price (per second)Free quota(Note)
fun-music-v1$0.000275No free quota

Speech synthesis (text-to-speech)

Qwen-TTS

International

When the deployment scope is International, model inference compute resources are dynamically scheduled globally (excluding mainland China). Static data is stored in the region that you select. The supported region for this deployment scope is Singapore.

Qwen3-TTS-Instruct-Flash

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Free quota (Note)
qwen3-tts-instruct-flash$0.11510,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-instruct-flash-2026-01-26$0.115

Qwen3-TTS-VD

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Free quota (Note)
qwen3-tts-vd-2026-01-26$0.11510,000 charactersValidity: 90 days after activating Model Studio

Qwen3-TTS-VC

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Free quota (Note)
qwen3-tts-vc-2026-01-22$0.11510,000 charactersValidity: 90 days after activating Model Studio

Qwen3-TTS-Flash

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Free quota (Note)
qwen3-tts-flash ** Currently equivalent to qwen3-tts-flash-2025-11-27$0.110,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-flash-2025-11-27$0.1
qwen3-tts-flash-2025-09-18$0.1For activations before 00:00, November 13, 2025: 2,000 charactersFor activations at or after 00:00, November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio

Chinese mainland

When the deployment scope is mainland China, model inference computing resources are limited to mainland China, and static data is stored in the region that you select. The region supported by this deployment scope is: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Qwen3-TTS-Instruct-Flash

You are billed based on the number of characters in the input text; the output is not charged.

Model name**Input price (per 10,000 characters)Output price (per 10,000 characters)
qwen3-tts-instruct-flash$0.115Not billed
qwen3-tts-instruct-flash-2026-01-26$0.115Not billed

Qwen3-TTS-VD

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Output price (per 10,000 characters)
qwen3-tts-vd-2026-01-26$0.115Not billed

Qwen3-TTS-VC

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Output price (per 10,000 characters)
qwen3-tts-vc-2026-01-22$0.115Not billed

Qwen3-TTS-Flash

You are billed based on the number of characters in the input text; the output is not charged.

Model nameInput price (per 10,000 characters)Output price (per 10,000 characters)
qwen3-tts-flash ** Currently equivalent to qwen3-tts-flash-2025-11-27$0.114682Not billed
qwen3-tts-flash-2025-11-27$0.114682Not billed
qwen3-tts-flash-2025-09-18$0.114682Not billed

Qwen-TTS

You are billed based on the number of input and output tokens.

Model name**Input price (per million tokens)Output price (per million tokens)
qwen-tts-flash$0.23$1.434
qwen-tts-latest$0.23$1.434
qwen-tts-2025-05-22$0.23$1.434
qwen-tts-2025-04-10$0.23$1.434

Qwen-TTS-Realtime

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the region you select. The supported region is Singapore.

Qwen3-TTS-Instruct-Flash-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersFree quota (note)
qwen3-tts-instruct-flash-realtime$0.14310,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-instruct-flash-realtime-2026-01-22$0.14310,000 charactersValidity: 90 days after activating Model Studio

Qwen3-TTS-VD-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersFree quota (note)
qwen3-tts-vd-realtime-2026-01-15$0.14335310,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-vd-realtime-2025-12-16$0.14335310,000 charactersValidity: 90 days after activating Model Studio

Qwen3-TTS-VC-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersFree quota (note)
qwen3-tts-vc-realtime-2026-01-15$0.1310,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-vc-realtime-2025-11-27

Qwen3-TTS-Flash-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersFree quota (note)
qwen3-tts-flash-realtime$0.13For Model Studio activated before 00:00 on November 13, 2025: 2,000 charactersFor Model Studio activated at or after 00:00 on November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-flash-realtime-2025-11-27$0.1310,000 charactersValidity: 90 days after activating Model Studio
qwen3-tts-flash-realtime-2025-09-18$0.13For Model Studio activated before 00:00 on November 13, 2025: 2,000 charactersFor Model Studio activated at or after 00:00 on November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in the region you select. The supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Qwen3-TTS-Instruct-Flash-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersOutput price
qwen3-tts-instruct-flash-realtime$0.143Not charged
qwen3-tts-instruct-flash-realtime-2026-01-22$0.143Not charged

Qwen3-TTS-VD-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersOutput price
qwen3-tts-vd-realtime-2026-01-15$0.143353Not charged
qwen3-tts-vd-realtime-2025-12-16$0.143353Not charged

Qwen3-TTS-VC-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersOutput price
qwen3-tts-vc-realtime-2026-01-15$0.143353Not charged
qwen3-tts-vc-realtime-2025-11-27

Qwen3-TTS-Flash-Realtime

Billing rules: Billing is based on the number of characters in the input text. Output is not billed.

ModelInput price / 10,000 charactersOutput price
qwen3-tts-flash-realtime$0.143353Not charged
qwen3-tts-flash-realtime-2025-11-27$0.143353Not charged
qwen3-tts-flash-realtime-2025-09-18$0.143353Not charged

Qwen-TTS-Realtime

Billing rules: Billing is based on the number of input and output tokens.

ModelInput price / 1,000,000 tokensOutput price / 1,000,000 tokens
qwen-tts-realtime$0.345$1.721
qwen-tts-realtime-latest$0.345$1.721
qwen-tts-realtime-2025-07-15$0.345$1.721

Qwen-TTS voice cloning

Billing: You are charged for each new voice.

International

The International deployment scope dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in Singapore.

Model namePrice (per voice)Free quota (Note)
qwen-voice-enrollment$0.011,000 voices/account

Chinese mainland

The Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model namePrice (per voice)
qwen-voice-enrollment$0.01

Qwen-TTS voice design

You are charged for each new voice you create.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in Singapore.

ModelPrice (per voice)Free quota(Note)
qwen-voice-design$0.210 voices/account

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelPrice (per voice)
qwen-voice-design$0.2

CosyVoice

International

If you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and your static data is stored in your selected region. The only supported region is Singapore.

Billing is based on the number of input characters; output is not charged.

ModelInput price (per 10,000 characters)Free quota (Note)
cosyvoice-v3-plus$0.2610,000 charactersThe quota is valid for 90 days after you activate Model Studio.
cosyvoice-v3-flash$0.13

Chinese mainland

If you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland, and your static data is stored in your selected region. The only supported region is China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Billing is based on the number of input characters; output is not charged.

ModelInput price (per 10,000 characters)Free quota (Note)
cosyvoice-v3.5-plus$0.22No free quota
cosyvoice-v3.5-flash$0.116
cosyvoice-v3-plus$0.2867
cosyvoice-v3-flash$0.1434
cosyvoice-v2$0.2867

Speech recognition and speech translation

Qwen-LiveTranslate-Flash-Realtime

Billing rules: Billing is based on the number of input and output tokens. For details on how tokens are calculated for different modalities, see Billing.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.

ModelInput price (per 1,000,000 tokens)Output price (per 1,000,000 tokens)Free quota (Note)
Input: AudioInput: ImageOutput: TextOutput: Audio
qwen3-livetranslate-flash-realtime$10$1.3$10$381 million tokens for each modality Validity: 90 days after you activate Model Studio
qwen3-livetranslate-flash-realtime-2025-09-22$10$1.3$10$38

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput price (per 1,000,000 tokens)Output price (per 1,000,000 tokens)
Input: AudioInput: ImageOutput: TextOutput: Audio
qwen3-livetranslate-flash-realtime$9.175$1.147$9.175$34.405
qwen3-livetranslate-flash-realtime-2025-09-22$9.175$1.147$9.175$34.405

Qwen-ASR

Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.

ModelInput priceFree quota (Note)
qwen3-asr-flash-filetrans$0.000035/second36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio
qwen3-asr-flash-filetrans-2025-11-17
qwen3-asr-flash ** Currently equivalent to qwen3-asr-flash-2025-09-08
qwen3-asr-flash-2026-02-10
qwen3-asr-flash-2025-09-08

US

When you select the US deployment scope, model inference compute resources are restricted to the United States. Your static data is stored in your selected region. This scope is available in the following region: US (Virginia). Note

Models in the US deployment scope do not offer a free quota.

Model**Input price
qwen3-asr-flash-us$0.000035/second
qwen3-asr-flash-2025-09-08-us

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput price
qwen3-asr-flash-filetrans$0.000032/second
qwen3-asr-flash-filetrans-2025-11-17
qwen3-asr-flash ** Currently equivalent to qwen3-asr-flash-2025-09-08
qwen3-asr-flash-2026-02-10
qwen3-asr-flash-2025-09-08

Qwen-ASR-Realtime

Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.

Model**Input priceFree quota (Note)
qwen3-asr-flash-realtime$0.000090/second36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio
qwen3-asr-flash-realtime-2026-02-10$0.000090/second
qwen3-asr-flash-realtime-2025-10-27$0.000090/second

Chinese mainland

When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

ModelInput price
qwen3-asr-flash-realtime$0.000047/second
qwen3-asr-flash-realtime-2026-02-10
qwen3-asr-flash-realtime-2025-10-27

Fun-ASR

Audio file recognition

Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.

International

When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.

ModelInput priceFree quota** (Note)**
fun-asr$0.000035/second36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio
fun-asr-2025-11-07
fun-asr-2025-08-25
fun-asr-mtl
fun-asr-mtl-2025-08-25

Mainland China

If the deployment scope is Mainland China, model inference compute resources are available only in Mainland China. Static data is stored in the region that you select. This deployment scope is supported in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model NameInput Price
fun-asr$0.000032/second
fun-asr-2025-11-07
fun-asr-2025-08-25
fun-asr-mtl
fun-asr-mtl-2025-08-25

Real-time speech recognition

Billing rule: You are billed based on the duration of the input audio in seconds. The output is not billed.

International

When the deployment scope is set to International, model inference compute resources are dynamically scheduled worldwide, excluding mainland China. Static data is stored in your selected region. This deployment scope is supported in the Singapore region.

Model nameInput priceFree tier** (Note)**
fun-asr-realtime$0.00009/second36,000 seconds (10 hours)Valid for 90 days
fun-asr-realtime-2025-11-07

Mainland China

When the deployment scope is set to Mainland China, model inference compute resources are restricted to mainland China. Static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput price
fun-asr-realtime$0.000047/second
fun-asr-realtime-2026-02-28
fun-asr-realtime-2025-11-07
fun-asr-realtime-2025-09-15
fun-asr-flash-8k-realtime$0.000032/second
fun-asr-flash-8k-realtime-2026-01-28

Paraformer

Audio file recognition

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Billing rules: You are billed based on the duration of the input audio in seconds. The output is not billed.

ModelInput price
paraformer-v2$0.000012/second
paraformer-8k-v2

Real-time speech recognition

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Billing rules: You are billed based on the duration of the input audio in seconds. The output is not billed.

ModelInput priceFree tier(Note)
paraformer-realtime-v2$0.000035/secondNo free tier
paraformer-realtime-8k-v2

Text embedding

Billing is based on input tokens. Output is not charged.

International

If you select the International deployment scope, model inference runs on dynamically scheduled resources worldwide, excluding the chinese mainland. Static data is stored in the region you select. Supported region: Singapore.

Model nameInput price (per 1M tokens)Free quota (Note)
text-embedding-v4$0.071,000,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio
text-embedding-v3$0.07500,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio

Chinese mainland

If you select the chinese mainland deployment scope, model inference is restricted to the chinese mainland. Static data is stored in the region you select. Supported region: China (Beijing). Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput price (per 1M tokens)
text-embedding-v4$0.072

China (Hong Kong)

If you select the China (Hong Kong) deployment scope, model inference is restricted to China (Hong Kong). Static data is stored in the region you select. Supported region: China (Hong Kong). Note

Models in the China (Hong Kong) deployment scope do not offer a free quota.

Model nameInput price (per 1M tokens)Free quota (Note)
text-embedding-v4$0.071,000,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio

Multimodal embedding

Billing is based on input tokens. Output is not charged.

International

If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. The system stores static data in the region you select. Supported region: Singapore.

ModelInput priceFree quota (Note)
tongyi-embedding-vision-plus$0.091 million tokensValidity: 90 days after activating Model Studio
tongyi-embedding-vision-flashImage/Video: $0.03Text: $0.09

Chinese mainland

If you select the Chinese mainland deployment scope, the system restricts model inference compute resources to the Chinese mainland. The system stores static data in the region you select. Supported region: China (Beijing).

ModelInput priceFree quota (Note)
qwen3-vl-embeddingImage/Video: $0.258Text: $0.11 million tokensValidity: 90 days after activating Model Studio
multimodal-embedding-v1Free trialUnlimited tokens

Text rerank

Billing rule: You are charged for input tokens only.

International

For the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in the Singapore region.

Model nameInput priceFree quota (Note)
qwen3-rerank$0.11 million tokensValid for 90 days after you activate Model Studio.

Chinese mainland

For the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in the China (Beijing) region. Note

Models in the Chinese mainland deployment scope do not offer a free quota.

Model nameInput price
gte-rerank-v2$0.115

Domain-specific models

Intent recognition

Note

Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.

Model nameInput priceOutput priceFree quota(Note)
tongyi-intent-detect-v3$0.058$0.144No free quota

Role playing

You are charged for input tokens and output tokens.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. The service stores static data in the region you select. Supported region: Singapore.

Model nameInput priceOutput priceFree quota(Note)
qwen-plus-character ** Discounts apply for session cache.$0.5$1.41 million tokens eachValid for 90 days after you activate Model Studio
qwen-flash-character Discounts apply for session cache.$0.05$0.4
qwen-plus-character-ja$0.5$1.4

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. The service stores static data in the region you select. Supported region: China (Beijing).

Model name**Input priceOutput priceFree quota(Note)
qwen-plus-character Discounts apply for session cache.$0.115$0.287No free quota
qwen-flash-character Discounts apply for session cache.$0.034$0.203No free quota

Mirror of Alibaba Cloud Model Studio docs for reference and RAG. Not affiliated with Alibaba Cloud.