Appearance
Model inference pricing
Tiered pricing rules
Some Model Studio models use tiered pricing. The total number of input tokens in a single request determines the unit price. All tokens in the request are then billed at the unit price for that tier.
For example, a model may have two pricing tiers: 0 < tokens ≤ 32K and 32K < tokens ≤ 128K. If a request contains 100K tokens, it falls into the second pricing tier (32K < 100K ≤ 128K). In this case, all 100K tokens are billed at the unit price for that tier.
Text generation: Qwen
Qwen-Max
You are charged for input tokens and output tokens.
If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. If the model supports context cache, only input tokens receive a discount. These two discounts cannot apply simultaneously.
International
With the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in your selected region. Supported region: Singapore.
Valid for 90 days after you activate Model Studio
| Model name | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** | Free quota (Note) |
|---|---|---|---|---|---|
| qwen3.6-max-preview ** Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 128K | $1.3 | $7.8 | 1 million tokens for input and output eachValid for 90 days after you activate Model Studio |
| 128K < tokens ≤ 256K | $2 | $12 | |||
| qwen3-max Currently equivalent to qwen3-max-2026-01-23 Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $1.2 | $6 | |
| 32K < tokens ≤ 128K | $2.4 | $12 | |||
| 128K < tokens ≤ 256K | $3 | $15 | |||
| qwen3-max-2026-01-23 | thinking and non-thinking | 0 < tokens ≤ 32K | $1.2 | $6 | |
| 32K < tokens ≤ 128K | $2.4 | $12 | |||
| 128K < tokens ≤ 256K | $3 | $15 | |||
| qwen3-max-2025-09-23 | non-thinking only | 0 < tokens ≤ 32K | $1.2 | $6 | |
| 32K < tokens ≤ 128K | $2.4 | $12 | |||
| 128K < tokens ≤ 256K | $3 | $15 | |||
| qwen3-max-preview Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $1.2 | $6 | |
| 32K < tokens ≤ 128K | $2.4 | $12 | |||
| 128K < tokens ≤ 256K | $3 | $15 |
More models
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|---|---|
| qwen-max ** Currently equivalent to qwen-max-2025-01-25 50% discount for batch inference | non-thinking only | no tiered pricing | $1.6 | $6.4 | 1 million tokens for input and output each Valid for 90 days after you activate Model Studio |
| qwen-max-latest | non-thinking only | no tiered pricing | $1.6 | $6.4 | |
| qwen-max-2025-01-25 | non-thinking only | no tiered pricing | $1.6 | $6.4 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|---|
| qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 Discount for context caching | non-thinking only | 0 < tokens ≤ 32K | $0.359 | $1.434 |
| 32K < tokens ≤ 128K | $0.574 | $2.294 | ||
| 128K < tokens ≤ 256K | $1.004 | $4.014 | ||
| qwen3-max-2025-09-23 | non-thinking only | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.434 | $5.735 | ||
| 128K < tokens ≤ 256K | $2.151 | $8.602 | ||
| qwen3-max-preview Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.434 | $5.735 | ||
| 128K < tokens ≤ 256K | $2.151 | $8.602 |
Chinese mainland
With the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|---|
| qwen3.6-max-preview ** Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 128K | $1.238 | $7.426 |
| 128K < tokens ≤ 256K | $2.063 | $12.377 | ||
| qwen3-max Currently equivalent to qwen3-max-2026-01-23 50% discount for batch inference Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $0.359 | $1.434 |
| 32K < tokens ≤ 128K | $0.574 | $2.294 | ||
| 128K < tokens ≤ 256K | $1.004 | $4.014 | ||
| qwen3-max-2026-01-23 | thinking and non-thinking | 0 < tokens ≤ 32K | $0.359 | $1.434 |
| 32K < tokens ≤ 128K | $0.574 | $2.294 | ||
| 128K < tokens ≤ 256K | $1.004 | $4.014 | ||
| qwen3-max-2025-09-23 | non-thinking only | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.434 | $5.735 | ||
| 128K < tokens ≤ 256K | $2.151 | $8.602 | ||
| qwen3-max-preview Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.434 | $5.735 | ||
| 128K < tokens ≤ 256K | $2.151 | $8.602 |
More models
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|---|
| qwen-max ** Currently equivalent to qwen-max-2024-09-19 | non-thinking only | no tiered pricing | $0.345 | $1.377 |
| qwen-max-latest | non-thinking only | no tiered pricing | $0.345 | $1.377 |
| qwen-max-2025-01-25 | non-thinking only | no tiered pricing | $0.345 | $1.377 |
| qwen-max-2024-09-19 | non-thinking only | no tiered pricing | $2.868 | $8.602 |
China (Hong Kong)
With the China (Hong Kong) deployment scope, both model inference compute resources and your static data are located in the China (Hong Kong) region. Note
Models in the China (Hong Kong) deployment scope do not offer a free quota.
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|---|
| qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 Discount for context caching | thinking and non-thinking | 0Model name** | Mode | Input tokens |
| --- | --- | --- | --- | --- |
| qwen3-max ** Currently equivalent to qwen3-max-2026-01-23 50% discount for batch inference Discount for context caching | thinking and non-thinking | 0 < tokens ≤ 32K | $1.2 | $6 |
| 32K < tokens ≤ 128K | $2.4 | $12 | ||
| 128K < tokens ≤ 256K | $3 | $15 | ||
| qwen3-max-2026-01-23 | thinking and non-thinking | 0 < tokens ≤ 32K | $1.2 | $6 |
| 32K < tokens ≤ 128K | $2.4 | $12 | ||
| 128K < tokens ≤ 256K | $3 | $15 |
Qwen-Plus
You are charged for input tokens and output tokens.
International
When the International deployment scope is selected, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the Singapore region.
| Model** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota** (note)** | |
|---|---|---|---|---|---|
| Non-thinking mode | Thinking mode (CoT + response) | ||||
| qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-02 | 0Model name** | Input token range | Input price | Output price | |
| --- | --- | --- | --- | --- | |
| Non-thinking mode | Thinking mode | ||||
| qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-02 | 0 < token ≤ 256K | $0.276 | $1.651 | $1.651 | |
| 256K < token ≤ 1M | $1.101 | $6.602 | $6.602 | ||
| qwen3.6-plus-2026-04-02 | 0 < token ≤ 256K | $0.276 | $1.651 | $1.651 | |
| 256K < token ≤ 1M | $1.101 | $6.602 | $6.602 | ||
| qwen3.5-plus Currently equivalent to qwen3.5-plus-2026-02-15 | 0 < token ≤ 128K | $0.115 | $0.688 | $0.688 | |
| 128K < token ≤ 256K | $0.287 | $1.72 | $1.72 | ||
| 256K < token ≤ 1M | $0.573 | $3.44 | $3.44 | ||
| qwen3.5-plus-2026-02-15 | 0 < token ≤ 128K | $0.115 | $0.688 | $0.688 | |
| 128K < token ≤ 256K | $0.287 | $1.72 | $1.72 | ||
| 256K < token ≤ 1M | $0.573 | $3.44 | $3.44 | ||
| qwen-plus Currently equivalent to qwen-plus-2025-12-01 | 0 < token ≤ 128K | $0.115 | $0.287 | $1.147 | |
| 128K < token ≤ 256K | $0.345 | $2.868 | $3.441 | ||
| 256K < token ≤ 1M | $0.689 | $6.881 | $9.175 | ||
| qwen-plus-2025-12-01 | 0 < token ≤ 128K | $0.115 | $0.287 | $1.147 | |
| 128K < token ≤ 256K | $0.345 | $2.868 | $3.441 | ||
| 256K < token ≤ 1M | $0.689 | $6.881 | $9.175 | ||
| qwen-plus-2025-09-11 | 0 < token ≤ 128K | $0.115 | $0.287 | $1.147 | |
| 128K < token ≤ 256K | $0.345 | $2.868 | $3.441 | ||
| 256K < token ≤ 1M | $0.689 | $6.881 | $9.175 | ||
| qwen-plus-2025-07-28 | 0 < token ≤ 128K | $0.115 | $0.287 | $1.147 | |
| 128K < token ≤ 256K | $0.345 | $2.868 | $3.441 | ||
| 256K < token ≤ 1M | $0.689 | $6.881 | $9.175 |
US
If you select the US deployment scope, model inference compute resources are located in the US. Static data is stored in your selected region. The supported region is US (Virginia). Note
Models in the US deployment scope do not offer a free quota.
| Model** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode (CoT + response) | |||
| qwen-plus-us | 0 < token ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < token ≤ 1M | $1.2 | $3.6 | $12 | |
| qwen-plus-2025-12-01-us | 0 < token ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < token ≤ 1M | $1.2 | $3.6 | $12 |
Chinese mainland
Selecting the Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought and response) | |||
| qwen3.6-plus ** Currently equivalent to qwen3.6-plus-2026-04-02 | 0Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
| --- | --- | --- | --- | |
| qwen-plus-2025-01-25 | no tiered pricing | $0.115 | $0.287 | |
| qwen-plus-2025-01-12 | no tiered pricing | $0.115 | $0.287 | |
| qwen-plus-2024-12-20 | no tiered pricing | $0.115 | $0.287 |
China (Hong Kong)
Selecting the China (Hong Kong) deployment scope restricts model inference compute resources to China (Hong Kong). Static data is also stored in this region. Note
Models in the China (Hong Kong) deployment scope do not offer a free quota.
| Model | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode | |||
| qwen-plus ** Currently equivalent to qwen-plus-2025-12-01 | 0 < tokens ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < tokens ≤ 1M | $1.2 | $3.6 | $12 | |
| qwen-plus-2025-12-01 | 0 < tokens ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < tokens ≤ 1M | $1.2 | $3.6 | $12 |
European Union
For deployments in the European Union, compute resources for model inference remain within the European Union. Static data is stored in your selected region. Supported region: Germany (Frankfurt). Note
Models in the EU deployment scope do not offer a free quota.
| Model** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking | Thinking (Chain of Thought (CoT) + response) | |||
| qwen-plus ** Currently equivalent to qwen-plus-2025-12-01 | 0 < token ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < token ≤ 1M | $1.2 | $3.6 | $12 | |
| qwen-plus-2025-12-01 | 0 < token ≤ 256K | $0.4 | $1.2 | $4 |
| 256K < token ≤ 1M | $1.2 | $3.6 | $12 |
Qwen-Flash
You are charged for input tokens and output tokens.
If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price. If the model supports context cache, only input tokens receive a discount. These two discounts cannot apply simultaneously.
International
When you select the International deployment scope, inference computation is dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is supported in the Singapore region.
| Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|---|
| qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-16 50% off for batch inference Discount for context caching | 0 < tokens ≤ 256K | $0.25 | $1.5 | 1 million input tokens and 1 million output tokens Valid for 90 days after you activate Model Studio. |
| 256K < tokens ≤ 1M | $1 | $4 | ||
| qwen3.6-flash-2026-04-16 | 0 < tokens ≤ 256K | $0.25 | $1.5 | |
| 256K < tokens ≤ 1M | $1 | $4 | ||
| qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-23 50% off for batch inference Discount for context caching | 0 < tokens ≤ 1M | $0.1 | $0.4 | |
| qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 1M | $0.1 | $0.4 | |
| qwen-flash Currently equivalent to qwen-flash-2025-07-28 50% off for batch inference Discount for context caching | 0 < tokens ≤ 256K | $0.05 | $0.4 | |
| 256K < tokens ≤ 1M | $0.25 | $2 | ||
| qwen-flash-2025-07-28 | 0 < tokens ≤ 256K | $0.05 | $0.4 | |
| 256K < tokens ≤ 1M | $0.25 | $2 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|
| qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-16 | 0 < tokens ≤ 256K | $0.165 | $0.99 |
| 256K < tokens ≤ 1M | $0.66 | $3.961 | |
| qwen3.6-flash-2026-04-16 | 0 < tokens ≤ 256K | $0.165 | $0.99 |
| 256K < tokens ≤ 1M | $0.66 | $3.961 | |
| qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 128K | $0.029 | $0.287 |
| 128K < tokens ≤ 256K | $0.115 | $1.147 | |
| 256K < tokens ≤ 1M | $0.172 | $1.72 | |
| qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 128K | $0.029 | $0.287 |
| 128K < tokens ≤ 256K | $0.115 | $1.147 | |
| 256K < tokens ≤ 1M | $0.172 | $1.72 | |
| qwen-flash Currently equivalent to qwen-flash-2025-07-28 Discount for context caching | 0 < tokens ≤ 128K | $0.022 | $0.216 |
| 128K < tokens ≤ 256K | $0.087 | $0.861 | |
| 256K < tokens ≤ 1M | $0.173 | $1.721 | |
| qwen-flash-2025-07-28 | 0 < tokens ≤ 128K | $0.022 | $0.216 |
| 128K < tokens ≤ 256K | $0.087 | $0.861 | |
| 256K < tokens ≤ 1M | $0.173 | $1.721 |
US
When you select the US deployment scope, inference computation is restricted to the US. Static data is stored in your selected region. This deployment scope is supported in the US (Virginia) region. Note
Models in the US deployment scope do not offer a free quota.
| Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|
| qwen-flash-us | 0 < tokens ≤ 256K | $0.05 | $0.4 |
| 256K < tokens ≤ 1M | $0.25 | $2 | |
| qwen-flash-2025-07-28-us | 0 < tokens ≤ 256K | $0.05 | $0.4 |
| 256K < tokens ≤ 1M | $0.25 | $2 |
Chinese mainland
When you select the Chinese mainland deployment scope, inference computation is restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|
| qwen3.6-flash ** Currently equivalent to qwen3.6-flash-2026-04-16 50% off for batch inference Discount for context caching | 0 < tokens ≤ 256K | $0.165 | $0.99 |
| 256K < tokens ≤ 1M | $0.66 | $3.961 | |
| qwen3.6-flash-2026-04-16 | 0 < tokens ≤ 256K | $0.165 | $0.99 |
| 256K < tokens ≤ 1M | $0.66 | $3.961 | |
| qwen3.5-flash Currently equivalent to qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 128K | $0.029 | $0.287 |
| 128K < tokens ≤ 256K | $0.115 | $1.147 | |
| 256K < tokens ≤ 1M | $0.172 | $1.72 | |
| qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 128K | $0.029 | $0.287 |
| 128K < tokens ≤ 256K | $0.115 | $1.147 | |
| 256K < tokens ≤ 1M | $0.172 | $1.72 | |
| qwen-flash Currently equivalent to qwen-flash-2025-07-28 Discount for context caching | 0 < tokens ≤ 128K | $0.022 | $0.216 |
| 128K < tokens ≤ 256K | $0.087 | $0.861 | |
| 256K < tokens ≤ 1M | $0.173 | $1.721 | |
| qwen-flash-2025-07-28 | 0 < tokens ≤ 128K | $0.022 | $0.216 |
| 128K < tokens ≤ 256K | $0.087 | $0.861 | |
| 256K < tokens ≤ 1M | $0.173 | $1.721 |
China (Hong Kong)
When you select the China (Hong Kong) deployment scope, inference computation is restricted to China (Hong Kong). Static data is stored in your selected region. This deployment scope is supported in the China (Hong Kong) region. Note
Models in the China (Hong Kong) deployment scope do not offer a free quota.
| Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|
| qwen3.5-flash ** Currently equivalent to qwen3.5-flash-2026-02-23 Discount for context caching | 0 < tokens ≤ 1M | $0.1 | $0.4 |
| qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 1M | $0.1 | $0.4 |
European Union
When you select the European Union deployment scope, inference computation is restricted to the European Union. Static data is stored in your selected region. This deployment scope is supported in the Germany (Frankfurt) region. Note
Models in the EU deployment scope do not offer a free quota.
| Model name** | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|---|
| qwen3.5-flash ** Currently equivalent to qwen3.5-flash-2026-02-23 Discount for context caching | 0 < tokens ≤ 1M | $0.1 | $0.4 |
| qwen3.5-flash-2026-02-23 | 0 < tokens ≤ 1M | $0.1 | $0.4 |
Qwen-Turbo
Note
Qwen-Turbo will no longer be updated. We recommend using Qwen-Flash instead. You are charged for input tokens and output tokens.
If the model supports batch calls, the unit price for both input and output tokens is 50% of the real-time inference price.
International
In the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought + response) | |||
| qwen-turbo ** Currently equivalent to qwen-turbo-2025-04-28 50% discount for batch inference | $0.05 | $0.2 | $0.5 | 1 million tokens each Valid for 90 days after activating Model Studio |
| qwen-turbo-latest | $0.05 | $0.2 | $0.5 | |
| qwen-turbo-2025-04-28 | $0.05 | $0.2 | $0.5 |
More models
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|
| qwen-turbo-2024-11-01 | $0.05 | $0.2 | 1 million tokens each Valid for 90 days after activating Model Studio |
Chinese mainland
In the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought + response) | ||
| qwen-turbo ** Currently equivalent to qwen-turbo-2025-04-28 | $0.044 | $0.087 | $0.431 |
| qwen-turbo-latest | $0.044 | $0.087 | $0.431 |
| qwen-turbo-2025-07-15 | $0.044 | $0.087 | $0.431 |
| qwen-turbo-2025-04-28 | $0.044 | $0.087 | $0.431 |
QwQ
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|
| qwq-plus ** Currently equivalent to qwq-plus-2025-03-05 | $0.8 | $2.4 | 1 million tokens Valid for 90 days after you activate Model Studio |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are confined to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwq-plus ** Currently equivalent to qwq-plus-2025-03-05 | $0.230 | $0.574 |
| qwq-plus-latest | $0.230 | $0.574 |
| qwq-plus-2025-03-05 | $0.230 | $0.574 |
Qwen-Long
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|
| ** **qwen-long-latest | $0.072 | $0.287 | No free quota |
| qwen-long-2025-01-25 | $0.072 | $0.287 |
Qwen-Omni
Billing is based on the number of input and output tokens. For token calculation rules for different modalities, see Billing and Throttling.
International
If you select the International deployment scope, model inference resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.
| Model | Input price | Output price | Free quota (Note) | ||
|---|---|---|---|---|---|
| Text/image/video | Audio | Text ** For multimodal input | Text+audio** ** Audio billed only | ||
| qwen3.5-omni-plus Currently equivalent to qwen3.5-omni-plus-2026-03-15 | $1.4 | $11 | $8.3 | $44 | 1 million tokensThis quota expires 90 days after you activate Model Studio. |
| qwen3.5-omni-plus-2026-03-15 | $1.4 | $11 | $8.3 | $44 | |
| qwen3.5-omni-flash Currently equivalent to qwen3.5-omni-flash-2026-03-15 | $0.4 | $3 | $2.2 | $11.9 | |
| qwen3.5-omni-flash-2026-03-15 | $0.4 | $3 | $2.2 | $11.9 |
More models
| Model** | Mode | Input price | Output price | Free quota (Note) | ||||
|---|---|---|---|---|---|---|---|---|
| Text | Audio | Image/video | Text ** For text-only input | Text** ** For multimodal input | Text+audio** ** Audio billed only | |||
| qwen3-omni-flash Currently equivalent to qwen3-omni-flash-2025-12-01 | non-thinking and thinking modes | $0.43 | $3.81 | $0.78 | $1.66 | $3.06 | $15.11 | 1 million tokens (modality-agnostic)This quota expires 90 days after you activate Model Studio. |
| qwen3-omni-flash-2025-12-01 | non-thinking and thinking modes | $0.43 | $3.81 | $0.78 | $1.66 | $3.06 | $15.11 | |
| qwen3-omni-flash-2025-09-15 | non-thinking and thinking modes | $0.43 | $3.81 | $0.78 | $1.66 | $3.06 | $15.11 | |
| qwen-omni-turbo Currently equivalent to qwen-omni-turbo-2025-03-26 | non-thinking mode | $0.07 | $4.44 | $0.21 | $0.27 | $0.63 | $8.89 | |
| qwen-omni-turbo-latest | non-thinking mode | $0.07 | $4.44 | $0.21 | $0.27 | $0.63 | $8.89 | |
| qwen-omni-turbo-2025-03-26 | non-thinking mode | $0.07 | $4.44 | $0.21 | $0.27 | $0.63 | $8.89 |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Input price | Output price | ||
|---|---|---|---|---|
| Text/image/video | Audio | Text ** For multimodal input | Text+audio** ** Audio billed only | |
| qwen3.5-omni-plus Currently equivalent to qwen3.5-omni-plus-2026-03-15 | $0.96 | $7.29 | $5.5 | $29.29 |
| qwen3.5-omni-plus-2026-03-15 | $0.96 | $7.29 | $5.5 | $29.29 |
| qwen3.5-omni-flash Currently equivalent to qwen3.5-omni-flash-2026-03-15 | $0.3 | $2.48 | $1.83 | $9.9 |
| qwen3.5-omni-flash-2026-03-15 | $0.3 | $2.48 | $1.83 | $9.9 |
More models
| Model** | Mode | Input price | Output price | ||||
|---|---|---|---|---|---|---|---|
| Text | Audio | Image/video | Text ** For text-only input | Text** ** For multimodal input | Text+audio** ** Audio billed only | ||
| qwen3-omni-flash Currently equivalent to qwen3-omni-flash-2025-12-01 | non-thinking and thinking modes | $0.258 | $2.265 | $0.473 | $0.989 | $1.821 | $8.974 |
| qwen3-omni-flash-2025-12-01 | non-thinking and thinking modes | $0.258 | $2.265 | $0.473 | $0.989 | $1.821 | $8.974 |
| qwen3-omni-flash-2025-09-15 | non-thinking and thinking modes | $0.258 | $2.265 | $0.473 | $0.989 | $1.821 | $8.974 |
| qwen-omni-turbo Currently equivalent to qwen-omni-turbo-2025-03-26 | non-thinking mode | $0.058 | $3.584 | $0.216 | $0.230 | $0.646 | $7.168 |
| qwen-omni-turbo-latest | non-thinking mode | $0.058 | $3.584 | $0.216 | $0.230 | $0.646 | $7.168 |
| qwen-omni-turbo-2025-03-26 | non-thinking mode | $0.058 | $3.584 | $0.216 | $0.230 | $0.646 | $7.168 |
| qwen-omni-turbo-2025-01-19 | non-thinking mode | $0.058 | $3.584 | $0.216 | $0.230 | $0.646 | $7.168 |
Qwen-Omni-Realtime
Billing is based on input and output tokens. See billing and rate limits for how tokens are calculated for different modalities.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in your selected region. The supported region for this deployment scope is Asia Pacific SE 1 (Singapore).
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) | ||
|---|---|---|---|---|---|
| Text and image | Audio | Text ** Multimodal input | Text + audio** ** Billed for audio only | ||
| qwen3.5-omni-plus-realtime | $2.1 | $16.5 | $12.4 | $62 | 1 million input tokens and 1 million output tokensValid for 90 days after activating Model Studio. |
| qwen3.5-omni-plus-realtime-2026-03-15 | $2.1 | $16.5 | $12.4 | $62 | |
| qwen3.5-omni-flash-realtime | $0.55 | $4.5 | $3.3 | $17.7 | |
| qwen3.5-omni-flash-realtime-2026-03-15 | $0.55 | $4.5 | $3.3 | $17.7 |
More models
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) | ||||
|---|---|---|---|---|---|---|---|
| Text | Audio | Image | Text ** For text-only input | Text** ** Multimodal input | Text + audio** ** Billed for audio only | ||
| qwen3-omni-flash-realtime | $0.52 | $4.57 | $0.94 | $1.99 | $3.67 | $18.13 | 1 million input tokens and 1 million output tokens (modality-agnostic)Valid for 90 days after activating Model Studio. |
| qwen3-omni-flash-realtime-2025-12-01 | $0.52 | $4.57 | $0.94 | $1.99 | $3.67 | $18.13 | |
| qwen3-omni-flash-realtime-2025-09-15 | $0.52 | $4.57 | $0.94 | $1.99 | $3.67 | $18.13 | |
| qwen-omni-turbo-realtime | $0.270 | $4.440 | $0.840 | $1.070 | $2.520 | $8.890 | |
| qwen-omni-turbo-realtime-latest | $0.270 | $4.440 | $0.840 | $1.070 | $2.520 | $8.890 | |
| qwen-omni-turbo-realtime-2025-05-08 | $0.270 | $4.440 | $0.840 | $1.070 | $2.520 | $8.890 |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | ||
|---|---|---|---|---|
| Text and image | Audio | Text ** Multimodal input | Text + audio** ** Billed for audio only | |
| qwen3.5-omni-plus-realtime | $1.38 | $11 | $8.25 | $41.26 |
| qwen3.5-omni-plus-realtime-2026-03-15 | $1.38 | $11 | $8.25 | $41.26 |
| qwen3.5-omni-flash-realtime | $0.45 | $3.71 | $2.75 | $14.71 |
| qwen3.5-omni-flash-realtime-2026-03-15 | $0.45 | $3.71 | $2.75 | $14.71 |
More models
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | ||||
|---|---|---|---|---|---|---|
| Text | Audio | Image | Text ** For text-only input | Text** ** Multimodal input | Text + audio** ** Billed for audio only | |
| qwen3-omni-flash-realtime | $0.315 | $2.709 | $0.559 | $1.19 | $2.179 | $10.766 |
| qwen3-omni-flash-realtime-2025-12-01 | $0.315 | $2.709 | $0.559 | $1.19 | $2.179 | $10.766 |
| qwen3-omni-flash-realtime-2025-09-15 | $0.315 | $2.709 | $0.559 | $1.19 | $2.179 | $10.766 |
| qwen-omni-turbo-realtime | $0.230 | $3.584 | $0.861 | $0.918 | $2.581 | $7.168 |
| qwen-omni-turbo-realtime-latest | $0.230 | $3.584 | $0.861 | $0.918 | $2.581 | $7.168 |
| qwen-omni-turbo-realtime-2025-05-08 | $0.230 | $3.584 | $0.861 | $0.918 | $2.581 | $7.168 |
QVQ
Billing is based on the number of input and output tokens. See Billing and rate limits for details on how tokens are calculated for different modalities.
International
If you select the International deployment scope, model inference resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.
| Model** | Input price | Output price | Free quota(Note) |
|---|---|---|---|
| qvq-max ** Currently equivalent to qvq-max-2025-03-25 | $1.2 | $4.8 | 1 million input and 1 million output tokens Valid for 90 days after activating Model Studio |
| qvq-max-latest | $1.2 | $4.8 | |
| qvq-max-2025-03-25 | $1.2 | $4.8 |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Input price | Output price |
|---|---|---|
| qvq-max ** Currently equivalent to qvq-max-2025-03-25 | $1.147 | $4.588 |
| qvq-max-latest | $1.147 | $4.588 |
| qvq-max-2025-05-15 | $1.147 | $4.588 |
| qvq-max-2025-03-25 | $1.147 | $4.588 |
| qvq-plus Currently equivalent to qvq-plus-2025-05-15 | $0.287 | $0.717 |
| qvq-plus-latest | $0.287 | $0.717 |
| qvq-plus-2025-05-15 | $0.287 | $0.717 |
Qwen-VL
You are charged for input tokens and output tokens.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in Singapore, the only supported region for this scope.
| Model name** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** Chain of Thought (CoT) + response** | Free quota (Note) |
|---|---|---|---|---|---|
| qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.2 | $1.6 | 1 million tokens each for input and output Valid for 90 days after you activate Model Studio. |
| 32K < tokens ≤ 128K | $0.3 | $2.4 | |||
| 128K < tokens ≤ 256K | $0.6 | $4.8 | |||
| qwen3-vl-plus-2025-12-19 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.2 | $1.6 | |
| 32K < tokens ≤ 128K | $0.3 | $2.4 | |||
| 128K < tokens ≤ 256K | $0.6 | $4.8 | |||
| qwen3-vl-plus-2025-09-23 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.2 | $1.6 | |
| 32K < tokens ≤ 128K | $0.3 | $2.4 | |||
| 128K < tokens ≤ 256K | $0.6 | $4.8 | |||
| qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discount for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 | |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | |||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | |||
| qwen3-vl-flash-2026-01-22 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 | |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | |||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | |||
| qwen3-vl-flash-2025-10-15 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 | |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | |||
| 128K < tokens ≤ 256K | $0.12 | $0.96 |
More models
| Model name** | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|---|
| qwen-vl-max ** Currently equivalent to qwen-vl-max-2025-08-13 Discount for context caching | No tiered pricing | $0.8 | $3.2 | 1 million tokens each for input and outputValid for 90 days after you activate Model Studio. |
| qwen-vl-max-latest | No tiered pricing | $0.8 | $3.2 | |
| qwen-vl-max-2025-08-13 | No tiered pricing | $0.8 | $3.2 | |
| qwen-vl-max-2025-04-08 | No tiered pricing | $0.8 | $3.2 | |
| qwen-vl-plus Currently equivalent to qwen-vl-plus-2025-08-15 Discount for context caching | No tiered pricing | $0.21 | $0.63 | |
| qwen-vl-plus-latest | No tiered pricing | $0.21 | $0.63 | |
| qwen-vl-plus-2025-08-15 | No tiered pricing | $0.21 | $0.63 | |
| qwen-vl-plus-2025-05-07 | No tiered pricing | $0.21 | $0.63 | |
| qwen-vl-plus-2025-01-25 | No tiered pricing | $0.21 | $0.63 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model** | Mode | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|---|
| qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.143 | $1.434 |
| 32K < tokens ≤ 128K | $0.215 | $2.15 | ||
| 128K < tokens ≤ 256K | $0.43 | $4.301 | ||
| qwen3-vl-plus-2025-09-23 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.143 | $1.434 |
| 32K < tokens ≤ 128K | $0.215 | $2.15 | ||
| 128K < tokens ≤ 256K | $0.43 | $4.301 | ||
| qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2025-10-15 Discount for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.022 | $0.215 |
| 32K < tokens ≤ 128K | $0.043 | $0.43 | ||
| 128K < tokens ≤ 256K | $0.086 | $0.859 | ||
| qwen3-vl-flash-2025-10-15 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.022 | $0.215 |
| 32K < tokens ≤ 128K | $0.043 | $0.43 | ||
| 128K < tokens ≤ 256K | $0.086 | $0.859 |
US
When the US deployment scope is selected, model inference resources are restricted to the US. Static data is stored in your selected region. The supported region is US (Virginia). Note
Models in the US deployment scope do not offer a free quota.
| Model name** | Mode | Input tokens per request | Input price (per 1M tokens) | Output price (per 1M tokens) ** Chain of thought + Response** |
|---|---|---|---|---|
| qwen3-vl-flash-us ** Discounts available for context caching | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | ||
| qwen3-vl-flash-2026-01-22-us | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | ||
| qwen3-vl-flash-2025-10-15-us | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 |
Chinese mainland
When you select the Chinese mainland as the deployment scope, model inference compute resources are restricted to the Chinese mainland, and static data is stored in your selected region. China (Beijing) is the only supported region for this deployment scope. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Mode | Input tokens | Input price | Output price ** CoT and response** |
|---|---|---|---|---|
| qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discounts are available for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.143 | $1.434 |
| 32K < tokens ≤ 128K | $0.215 | $2.15 | ||
| 128K < tokens ≤ 256K | $0.43 | $4.301 | ||
| qwen3-vl-plus-2025-12-19 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.143 | $1.434 |
| 32K < tokens ≤ 128K | $0.215 | $2.15 | ||
| 128K < tokens ≤ 256K | $0.43 | $4.301 | ||
| qwen3-vl-plus-2025-09-23 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.143 | $1.434 |
| 32K < tokens ≤ 128K | $0.215 | $2.15 | ||
| 128K < tokens ≤ 256K | $0.43 | $4.301 | ||
| qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discounts are available for context caching | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.022 | $0.215 |
| 32K < tokens ≤ 128K | $0.043 | $0.43 | ||
| 128K < tokens ≤ 256K | $0.086 | $0.859 | ||
| qwen3-vl-flash-2026-01-22 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.022 | $0.215 |
| 32K < tokens ≤ 128K | $0.043 | $0.43 | ||
| 128K < tokens ≤ 256K | $0.086 | $0.859 | ||
| qwen3-vl-flash-2025-10-15 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.022 | $0.215 |
| 32K < tokens ≤ 128K | $0.043 | $0.43 | ||
| 128K < tokens ≤ 256K | $0.086 | $0.859 |
More models
| Model** | Input tokens | Input price | Output price |
|---|---|---|---|
| qwen-vl-max ** Currently equivalent to qwen-vl-max-2025-08-13 Discounts are available for context caching | No tiered pricing | $0.23 | $0.574 |
| qwen-vl-max-latest | No tiered pricing | $0.23 | $0.574 |
| qwen-vl-max-2025-08-13 | No tiered pricing | $0.23 | $0.574 |
| qwen-vl-max-2025-04-08 | No tiered pricing | $0.431 | $1.291 |
| qwen-vl-max-2025-04-02 | No tiered pricing | $0.431 | $1.291 |
| qwen-vl-max-2025-01-25 | No tiered pricing | $0.431 | $1.291 |
| qwen-vl-max-2024-12-30 | No tiered pricing | $0.431 | $1.291 |
| qwen-vl-max-2024-11-19 | No tiered pricing | $0.431 | $1.291 |
| qwen-vl-plus Currently equivalent to qwen-vl-plus-2025-08-15 Discounts are available for context caching | No tiered pricing | $0.115 | $0.287 |
| qwen-vl-plus-latest | No tiered pricing | $0.115 | $0.287 |
| qwen-vl-plus-2025-08-15 | No tiered pricing | $0.115 | $0.287 |
| qwen-vl-plus-2025-07-10 | No tiered pricing | $0.022 | $0.216 |
| qwen-vl-plus-2025-05-07 | No tiered pricing | $0.216 | $0.646 |
| qwen-vl-plus-2025-01-25 | No tiered pricing | $0.216 | $0.646 |
| qwen-vl-plus-2025-01-02 | No tiered pricing | $0.216 | $0.646 |
China (Hong Kong)
Selecting the China (Hong Kong) deployment scope places both model inference compute resources and static data in the China (Hong Kong) region. Note
Models in the China (Hong Kong) deployment scope do not offer a free quota.
| Model** | Mode | Input tokens | Input price | Output price ** CoT + answer** |
|---|---|---|---|---|
| qwen3-vl-plus ** Currently equivalent to qwen3-vl-plus-2025-12-19 Discount for context caching | non-thought and thought modes | 0Model** | Mode | Input tokens |
| --- | --- | --- | --- | --- |
| qwen3-vl-plus ** Discounts available for context caching | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.2 | $1.6 |
| 32K < tokens ≤ 128K | $0.3 | $2.4 | ||
| 128K < tokens ≤ 256K | $0.6 | $4.8 | ||
| qwen3-vl-flash Currently equivalent to qwen3-vl-flash-2026-01-22 Discounts available for context caching | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | ||
| qwen3-vl-flash-2026-01-22 | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 | ||
| qwen3-vl-flash-2025-10-15 | Thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.05 | $0.4 |
| 32K < tokens ≤ 128K | $0.075 | $0.6 | ||
| 128K < tokens ≤ 256K | $0.12 | $0.96 |
Qwen-OCR
You are charged for input tokens and output tokens.
International
Selecting the International deployment scope dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.
| Model** | Input price | Output price | Free quota (see note) |
|---|---|---|---|
| qwen-vl-ocr | $0.07 | $0.16 | 1 million tokens for each model This quota expires 90 days after you activate Model Studio. |
| qwen-vl-ocr-2025-11-20 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model | Input price | Output price |
|---|---|---|
| qwen-vl-ocr | $0.043 | $0.072 |
| qwen-vl-ocr-2025-11-20 | $0.043 | $0.072 |
Chinese mainland
Selecting the Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price | Output price |
|---|---|---|
| qwen-vl-ocr | $0.043 | $0.072 |
| qwen-vl-ocr-latest | $0.043 | $0.072 |
| qwen-vl-ocr-2025-11-20 | $0.043 | $0.072 |
| qwen-vl-ocr-2025-08-28 | $0.717 | $0.717 |
| qwen-vl-ocr-2025-04-13 | $0.717 | $0.717 |
| qwen-vl-ocr-2024-10-28 | $0.717 | $0.717 |
Qwen-Math
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model name | Input price | Output price | Free quota**(note)** |
|---|---|---|---|
| qwen-math-plus | $0.574 | $1.721 | No free quota |
| qwen-math-plus-latest | $0.574 | $1.721 | |
| qwen-math-plus-2024-09-19 | $0.574 | $1.721 | |
| qwen-math-plus-2024-08-16 | $0.574 | $1.721 | |
| qwen-math-turbo | $0.287 | $0.861 | |
| qwen-math-turbo-latest | $0.287 | $0.861 | |
| qwen-math-turbo-2024-09-19 | $0.287 | $0.861 |
Qwen-Coder
You are charged for input tokens and output tokens.
If the model supports context cache, only input tokens receive a discount.
International
In the International deployment scope, model inference runs on compute resources that are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in your selected region. The supported region for this scope is Singapore.
| Model | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(note) |
|---|---|---|---|---|
| qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount | 0Model** | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) |
| --- | --- | --- | --- | |
| qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount | 0Model** | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) |
| --- | --- | --- | --- | |
| qwen3-coder-plus ** Currently equivalent to qwen3-coder-plus-2025-09-23 Eligible for context caching discount | 0Model** | Input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) |
| --- | --- | --- | --- | |
| qwen-coder-plus ** Currently equivalent to qwen-coder-plus-2024-11-06 | no tiered pricing | $0.502 | $1.004 | |
| qwen-coder-plus-latest | no tiered pricing | $0.502 | $1.004 | |
| qwen-coder-plus-2024-11-06 | no tiered pricing | $0.502 | $1.004 | |
| qwen-coder-turbo Currently equivalent to qwen-coder-turbo-2024-09-19 | no tiered pricing | $0.287 | $0.861 | |
| qwen-coder-turbo-latest | no tiered pricing | $0.287 | $0.861 | |
| qwen-coder-turbo-2024-09-19 | no tiered pricing | $0.287 | $0.861 |
Qwen-MT
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the Singapore region.
| Model** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(note) |
|---|---|---|---|
| qwen-mt-plus | $2.46 | $7.37 | 1M input and 1M output tokens Valid for 90 days after you activate Model Studio |
| qwen-mt-flash | $0.16 | $0.49 | |
| qwen-mt-lite | $0.12 | $0.36 | |
| qwen-mt-turbo | $0.16 | $0.49 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwen-mt-plus | $0.259 | $0.775 |
| qwen-mt-flash | $0.101 | $0.280 |
| qwen-mt-lite | $0.086 | $0.229 |
US
If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data is stored in your selected region. This deployment scope is available in the US (Virginia) region. Note
Models in the US deployment scope do not offer a free quota.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwen-mt-lite-us | $0.12 | $0.36 |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwen-mt-plus | $0.259 | $0.775 |
| qwen-mt-flash | $0.101 | $0.280 |
| qwen-mt-lite | $0.086 | $0.229 |
| qwen-mt-turbo | $0.101 | $0.280 |
Qwen data mining
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model | Input price | Output price | Free quota (Note) |
|---|---|---|---|
| qwen-doc-turbo | $0.087 | $0.144 | No free quota |
Qwen deep research
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model name | Input price (per million tokens) | Output price (per million tokens) | **Free quota **(Note) |
|---|---|---|---|
| qwen-deep-research | $7.742 | $23.367 | None |
Text generation - Qwen - Open source
Qwen3.6
You are charged for input tokens and output tokens.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).
| Model | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought + response) | |||
| qwen3.6-35b-a3b | 0 < token ≤ 256K | $0.248 | $1.485 | $1.485 |
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland), while static data is stored in your selected region. Supported region: Singapore (Singapore).
| Model | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota**(Note)** | |
|---|---|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought + response) | ||||
| qwen3.6-35b-a3b | 0 < token ≤ 256K | $0.375 | $2.25 | $2.25 | 1 million tokens for each model. Validity: 90 days after you activate Model Studio. |
| qwen3.6-27b | 0 < token ≤ 256K | $0.6 | $3.6 | $3.6 |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland, while static data is stored in your selected region. Supported region: China (Beijing).
| Model | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode (chain of thought + response) | |||
| qwen3.6-35b-a3b | 0 < token ≤ 256K | $0.248 | $1.485 | $1.485 |
| qwen3.6-27b | 0 < token ≤ 256K | $0.412564 | $2.475384 | $2.475384 |
Qwen3.5
You are charged for input tokens and output tokens.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).
| Model name | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking | Thinking (CoT + response) | |||
| qwen3.5-397b-a17b | 0 < token ≤ 128K | $0.172 | $1.032 | $1.032 |
| 128K < token ≤ 256K | $0.43 | $2.58 | $2.58 | |
| qwen3.5-122b-a10b | 0 < token ≤ 128K | $0.115 | $0.917 | $0.917 |
| 128K < token ≤ 256K | $0.287 | $2.294 | $2.294 | |
| qwen3.5-27b | 0 < token ≤ 128K | $0.086 | $0.688 | $0.688 |
| 128K < token ≤ 256K | $0.258 | $2.064 | $2.064 | |
| qwen3.5-35b-a3b | 0 < token ≤ 128K | $0.057 | $0.459 | $0.459 |
| 128K < token ≤ 256K | $0.229 | $1.835 | $1.835 |
International
If you select the International deployment scope, the platform dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data remains in your selected region. Supported region: Singapore.
| Model name | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota** (note)** | |
|---|---|---|---|---|---|
| Non-thinking | Thinking (CoT + response) | ||||
| qwen3.5-397b-a17b | 0 < token ≤ 256K | $0.6 | $3.6 | $3.6 | 1 million tokens per model This quota is valid for 90 days after you activate Model Studio. |
| qwen3.5-122b-a10b | 0 < token ≤ 256K | $0.4 | $3.2 | $3.2 | |
| qwen3.5-27b | 0 < token ≤ 256K | $0.3 | $2.4 | $2.4 | |
| qwen3.5-35b-a3b | 0 < token ≤ 256K | $0.25 | $2 | $2 |
Chinese mainland
If you select the Chinese mainland deployment scope, the platform restricts model inference compute resources to the Chinese mainland. Static data remains in your selected region. Supported region: China (Beijing).
| Model name | Input token range | Input price (per 1M tokens) | Output price (per 1M tokens) | |
|---|---|---|---|---|
| Non-thinking | Thinking (CoT + response) | |||
| qwen3.5-397b-a17b | 0 < token ≤ 128K | $0.172 | $1.032 | $1.032 |
| 128K < token ≤ 256K | $0.43 | $2.58 | $2.58 | |
| qwen3.5-122b-a10b | 0 < token ≤ 128K | $0.115 | $0.917 | $0.917 |
| 128K < token ≤ 256K | $0.287 | $2.294 | $2.294 | |
| qwen3.5-27b | 0 < token ≤ 128K | $0.086 | $0.688 | $0.688 |
| 128K < token ≤ 256K | $0.258 | $2.064 | $2.064 | |
| qwen3.5-35b-a3b | 0 < token ≤ 128K | $0.057 | $0.459 | $0.459 |
| 128K < token ≤ 256K | $0.229 | $1.835 | $1.835 |
Qwen3
You are charged for input tokens and output tokens.
International
With the International deployment scope, model inference resources are dynamically scheduled worldwide (excluding the Chinese mainland), and static data is stored in your selected region. This scope is available in the Singapore region.
| Model | Mode | Input price | Output price | Free quota (note) | |
|---|---|---|---|---|---|
| Non-thinking mode | Thinking mode | ||||
| qwen3-next-80b-a3b-thinking | Thinking mode only | $0.15 | - | $1.2 | 1 million tokens per account. This quota is valid for 90 days after you activate Model Studio. |
| qwen3-next-80b-a3b-instruct | Non-thinking mode only | $0.15 | $1.2 | - | |
| qwen3-235b-a22b-thinking-2507 | Thinking mode only | $0.23 | - | $2.3 | |
| qwen3-235b-a22b-instruct-2507 | Non-thinking mode only | $0.23 | $0.92 | - | |
| qwen3-30b-a3b-thinking-2507 | Thinking mode only | $0.2 | - | $2.4 | |
| qwen3-30b-a3b-instruct-2507 | Non-thinking mode only | $0.2 | $0.8 | - | |
| qwen3-235b-a22b | Non-thinking & thinking modes | $0.7 | $2.8 | $8.4 | |
| qwen3-32b | Non-thinking & thinking modes | $0.16 | $0.64 | $0.64 | |
| qwen3-30b-a3b | Non-thinking & thinking modes | $0.2 | $0.8 | $2.4 | |
| qwen3-14b | Non-thinking & thinking modes | $0.35 | $1.4 | $4.2 | |
| qwen3-8b | Non-thinking & thinking modes | $0.18 | $0.7 | $2.1 | |
| qwen3-4b | Non-thinking & thinking modes | $0.11 | $0.42 | $1.26 | |
| qwen3-1.7b | Non-thinking & thinking modes | $0.11 | $0.42 | $1.26 | |
| qwen3-0.6b | Non-thinking & thinking modes | $0.11 | $0.42 | $1.26 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model | Mode | Input price | Output price | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode | |||
| qwen3-next-80b-a3b-thinking | Thinking mode only | $0.144 | - | $1.434 |
| qwen3-next-80b-a3b-instruct | Non-thinking mode only | $0.144 | $0.574 | - |
| qwen3-235b-a22b-thinking-2507 | Thinking mode only | $0.23 | - | $2.3 |
| qwen3-235b-a22b-instruct-2507 | Non-thinking mode only | $0.23 | $0.92 | - |
| qwen3-30b-a3b-thinking-2507 | Thinking mode only | $0.108 | - | $1.076 |
| qwen3-30b-a3b-instruct-2507 | Non-thinking mode only | $0.108 | $0.431 | - |
| qwen3-235b-a22b | Non-thinking & thinking modes | $0.287 | $1.147 | $2.868 |
| qwen3-32b | Non-thinking & thinking modes | $0.16 | $0.64 | $0.64 |
| qwen3-30b-a3b | Non-thinking & thinking modes | $0.108 | $0.431 | $1.076 |
| qwen3-14b | Non-thinking & thinking modes | $0.144 | $0.574 | $1.434 |
| qwen3-8b | Non-thinking & thinking modes | $0.072 | $0.287 | $0.717 |
Chinese mainland
With the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland, and static data is stored in your selected region. This scope is available in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Mode | Input price | Output price | |
|---|---|---|---|---|
| Non-thinking mode | Thinking mode | |||
| qwen3-next-80b-a3b-thinking | Thinking mode only | $0.144 | - | $1.434 |
| qwen3-next-80b-a3b-instruct | Non-thinking mode only | $0.144 | $0.574 | - |
| qwen3-235b-a22b-thinking-2507 | Thinking mode only | $0.287 | - | $2.868 |
| qwen3-235b-a22b-instruct-2507 | Non-thinking mode only | $0.287 | $1.147 | - |
| qwen3-30b-a3b-thinking-2507 | Thinking mode only | $0.108 | - | $1.076 |
| qwen3-30b-a3b-instruct-2507 | Non-thinking mode only | $0.108 | $0.431 | - |
| qwen3-235b-a22b | Non-thinking & thinking modes | $0.287 | $1.147 | $2.868 |
| qwen3-32b | Non-thinking & thinking modes | $0.287 | $1.147 | $2.868 |
| qwen3-30b-a3b | Non-thinking & thinking modes | $0.108 | $0.431 | $1.076 |
| qwen3-14b | Non-thinking & thinking modes | $0.144 | $0.574 | $1.434 |
| qwen3-8b | Non-thinking & thinking modes | $0.072 | $0.287 | $0.717 |
| qwen3-4b | Non-thinking & thinking modes | $0.044 | $0.173 | $0.431 |
| qwen3-1.7b | Non-thinking & thinking modes | $0.044 | $0.173 | $0.431 |
| qwen3-0.6b | Non-thinking & thinking modes | $0.044 | $0.173 | $0.431 |
QwQ - open source
You are charged for input tokens and output tokens.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|
| qwq-32b | $0.287 | $0.861 | No free quota |
QwQ-Preview
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|
| qwq-32b-preview | $0.287 | $0.861 | No free quota |
Qwen2.5
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, compute resources for model inference are dynamically scheduled globally, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.
| Model name | Input price (per million tokens) | Output price (per million tokens) | Free quota (note) |
|---|---|---|---|
| qwen2.5-14b-instruct-1m | $0.805 | $3.22 | 1 million tokens each for input and output Valid for 90 days after you activate Model Studio |
| qwen2.5-7b-instruct-1m | $0.368 | $1.47 | |
| qwen2.5-72b-instruct | $1.4 | $5.6 | |
| qwen2.5-32b-instruct | $0.7 | $2.8 | |
| qwen2.5-14b-instruct | $0.35 | $1.4 | |
| qwen2.5-7b-instruct | $0.175 | $0.7 |
Chinese mainland
If you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input price (per million tokens) | Output price (per million tokens) |
|---|---|---|
| qwen2.5-14b-instruct-1m | $0.144 | $0.431 |
| qwen2.5-7b-instruct-1m | $0.072 | $0.144 |
| qwen2.5-72b-instruct | $0.574 | $1.721 |
| qwen2.5-32b-instruct | $0.287 | $0.861 |
| qwen2.5-14b-instruct | $0.144 | $0.431 |
| qwen2.5-7b-instruct | $0.072 | $0.144 |
| qwen2.5-3b-instruct | $0.044 | $0.130 |
| qwen2.5-1.5b-instruct | Free for a limited time | |
| qwen2.5-0.5b-instruct |
QVQ
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|
| qvq-72b-preview | $1.721 | $5.161 | No free quota |
Qwen-Omni
You are billed for input and output tokens. For details on how tokens are counted for different modalities, see billing and rate limiting.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the Singapore region.
| Model | Input price | Output price | Free quota (Note) | ||||
|---|---|---|---|---|---|---|---|
| Text | Audio | Image/video | Text ** Text-only input | Text** ** Multimodal input | Text+audio** ** Billed for audio only | ||
| qwen2.5-omni-7b | $0.10 | $6.76 | $0.28 | $0.40 | $0.84 | $13.51 | 1 million tokens (modality-agnostic)Valid for 90 days after activating Model Studio. |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Input price | Output price | ||||
|---|---|---|---|---|---|---|
| Text | Audio | Image/video | Text ** Text-only input | Text** ** Multimodal input | Text+audio** ** Billed for audio only | |
| qwen2.5-omni-7b | $0.087 | $5.448 | $0.287 | $0.345 | $0.861 | $10.895 |
Qwen3-Omni-Captioner
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|
| qwen3-omni-30b-a3b-captioner | $3.81 | $3.06 | 1 million tokens Valid for 90 days after activating Model Studio. |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwen3-omni-30b-a3b-captioner | $2.265 | $1.821 |
Qwen-VL
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide, excluding the Chinese mainland. Static data resides in your selected region. The supported region is Singapore.
| Model name | Mode | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** | Free quota (note) |
|---|---|---|---|---|
| qwen3-vl-235b-a22b-thinking | thinking mode | $0.4 | $4 | 1 million tokens for each model. The quota is valid for 90 days after you activate Model Studio. |
| qwen3-vl-235b-a22b-instruct | non-thinking mode | $0.4 | $1.6 | |
| qwen3-vl-32b-thinking | thinking mode | $0.16 | $0.64 | |
| qwen3-vl-32b-instruct | non-thinking mode | $0.16 | $0.64 | |
| qwen3-vl-30b-a3b-thinking | thinking mode | $0.2 | $2.4 | |
| qwen3-vl-30b-a3b-instruct | non-thinking mode | $0.2 | $0.8 | |
| qwen3-vl-8b-thinking | thinking mode | $0.18 | $2.1 | |
| qwen3-vl-8b-instruct | non-thinking mode | $0.18 | $0.7 |
More models
| Model name | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (note) |
|---|---|---|---|
| qwen2.5-vl-72b-instruct | $2.8 | $8.4 | 1 million tokens for each model. The quota is valid for 90 days after you activate Model Studio. |
| qwen2.5-vl-32b-instruct | $1.4 | $4.2 | |
| qwen2.5-vl-7b-instruct | $0.35 | $1.05 | |
| qwen2.5-vl-3b-instruct | $0.21 | $0.63 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Mode | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|
| qwen3-vl-235b-a22b-thinking | thinking mode | $0.287 | $2.867 |
| qwen3-vl-235b-a22b-instruct | non-thinking mode | $0.287 | $1.147 |
| qwen3-vl-32b-thinking | thinking mode | $0.16 | $0.64 |
| qwen3-vl-32b-instruct | non-thinking mode | $0.16 | $0.64 |
| qwen3-vl-30b-a3b-thinking | thinking mode | $0.108 | $1.076 |
| qwen3-vl-30b-a3b-instruct | non-thinking mode | $0.108 | $0.431 |
| qwen3-vl-8b-thinking | thinking mode | $0.072 | $0.717 |
| qwen3-vl-8b-instruct | non-thinking mode | $0.072 | $0.287 |
Chinese mainland
If you select the Chinese mainland deployment scope, compute resources for model inference operate exclusively within the Chinese mainland. Static data resides in your selected region. The supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Mode | Input price (per 1M tokens) | Output price (per 1M tokens) ** CoT + response** |
|---|---|---|---|
| qwen3-vl-235b-a22b-thinking | thinking mode | $0.287 | $2.867 |
| qwen3-vl-235b-a22b-instruct | non-thinking mode | $0.287 | $1.147 |
| qwen3-vl-32b-thinking | thinking mode | $0.287 | $2.868 |
| qwen3-vl-32b-instruct | non-thinking mode | $0.287 | $1.147 |
| qwen3-vl-30b-a3b-thinking | thinking mode | $0.108 | $1.076 |
| qwen3-vl-30b-a3b-instruct | non-thinking mode | $0.108 | $0.431 |
| qwen3-vl-8b-thinking | thinking mode | $0.072 | $0.717 |
| qwen3-vl-8b-instruct | non-thinking mode | $0.072 | $0.287 |
More models
| Model name | Input price (per 1M tokens) | Output price (per 1M tokens) |
|---|---|---|
| qwen2.5-vl-72b-instruct | $2.294 | $6.881 |
| qwen2.5-vl-32b-instruct | $1.147 | $3.441 |
| qwen2.5-vl-7b-instruct | $0.287 | $0.717 |
| qwen2.5-vl-3b-instruct | $0.173 | $0.517 |
| qwen2-vl-72b-instruct | $2.294 | $6.881 |
| qwen2-vl-7b-instruct | Free for a limited time | |
| qwen2-vl-2b-instruct |
Qwen-Math
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model | Input price | Output price | Free quota (Note) |
|---|---|---|---|
| qwen2.5-math-72b-instruct | $0.574 | $1.721 | No free quota |
| qwen2.5-math-7b-instruct | $0.144 | $0.287 | |
| qwen2.5-math-1.5b-instruct | Free for a limited time |
Qwen-Coder
You are charged for input tokens and output tokens.
International
For the International deployment scope, model inference resources are scheduled dynamically worldwide (excluding the Chinese mainland). Your static data is stored in the selected region. The only supported region for this scope is Singapore.
| Model name | Input tokens | Input price (per million tokens) | Output price (per million tokens) | Free quota (note) |
|---|---|---|---|---|
| qwen3-coder-next | 0 < tokens ≤ 32K | $0.3 | $1.5 | 1 million input tokens and 1 million output tokens Validity: Within 90 days of activating Model Studio |
| 32K < tokens ≤ 128K | $0.5 | $2.5 | ||
| 128K < tokens ≤ 256K | $0.8 | $4 | ||
| qwen3-coder-480b-a35b-instruct | 0 < tokens ≤ 32K | $1.5 | $7.5 | |
| 32K < tokens ≤ 128K | $2.7 | $13.5 | ||
| 128K < tokens ≤ 200K | $4.5 | $22.5 | ||
| qwen3-coder-30b-a3b-instruct | 0 < tokens ≤ 32K | $0.45 | $2.25 | |
| 32K < tokens ≤ 128K | $0.75 | $3.75 | ||
| 128K < tokens ≤ 200K | $1.2 | $6 |
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Input tokens | Input price (per million tokens) | Output price (per million tokens) |
|---|---|---|---|
| qwen3-coder-480b-a35b-instruct | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.291 | $5.161 | |
| 128K < tokens ≤ 200K | $2.151 | $8.602 | |
| qwen3-coder-30b-a3b-instruct | 0 < tokens ≤ 32K | $0.216 | $0.861 |
| 32K < tokens ≤ 128K | $0.323 | $1.291 | |
| 128K < tokens ≤ 200K | $0.538 | $2.151 |
Chinese mainland
For the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland. Your static data is stored in the selected region. The only supported region for this scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input tokens | Input price (per million tokens) | Output price (per million tokens) |
|---|---|---|---|
| qwen3-coder-next | 0 < tokens ≤ 32K | $0.144 | $0.574 |
| 32K < tokens ≤ 128K | $0.216 | $0.861 | |
| 128K < tokens ≤ 256K | $0.359 | $1.434 | |
| qwen3-coder-480b-a35b-instruct | 0 < tokens ≤ 32K | $0.861 | $3.441 |
| 32K < tokens ≤ 128K | $1.291 | $5.161 | |
| 128K < tokens ≤ 200K | $2.151 | $8.602 | |
| qwen3-coder-30b-a3b-instruct | 0 < tokens ≤ 32K | $0.216 | $0.861 |
| 32K < tokens ≤ 128K | $0.323 | $1.291 | |
| 128K < tokens ≤ 200K | $0.538 | $2.151 | |
| qwen2.5-coder-32b-instruct | No tiered pricing | $0.287 | $0.861 |
| qwen2.5-coder-14b-instruct | No tiered pricing | $0.287 | $0.861 |
| qwen2.5-coder-7b-instruct | No tiered pricing | $0.144 | $0.287 |
| qwen2.5-coder-3b-instruct | No tiered pricing | free for a limited time | |
| qwen2.5-coder-1.5b-instruct | No tiered pricing | ||
| qwen2.5-coder-0.5b-instruct | No tiered pricing |
European Union
For the European Union deployment scope, model inference resources are restricted to the European Union. Your static data is stored in the selected region. For this scope, the only supported region is Germany (Frankfurt). Note
Models in the EU deployment scope do not offer a free quota.
| Model name | Input tokens | Input price (per million tokens) | Output price (per million tokens) |
|---|---|---|---|
| qwen3-coder-next | 0 < tokens ≤ 32K | $0.3 | $1.5 |
| 32K < tokens ≤ 128K | $0.5 | $2.5 | |
| 128K < tokens ≤ 256K | $0.8 | $4 |
Text generation - third-party models
DeepSeek
You are charged for input tokens and output tokens.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).
| Model name | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota |
|---|---|---|---|
| deepseek-v4-pro ** Discounts are available for context caching. | $1.65 | $3.3 | No free quota |
| deepseek-v4-flash Discounts are available for context caching. | $0.14 | $0.28 |
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland), while static data is stored in your selected region. The supported region for this deployment scope is Singapore.
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota |
|---|---|---|---|
| deepseek-v4-pro ** Discounts are available for context caching. | $2.400 | $4.800 | 1 million tokensValidity period: 90 days after activating Model Studio |
| deepseek-v4-flash Discounts are available for context caching. | $0.200 | $0.400 | 1 million tokensValidity period: 90 days after activating Model Studio |
| deepseek-v3.2 Discounts are available for context caching. | $0.57 | $1.71 | 1 million tokensValidity period: 90 days after activating Model Studio |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland, while static data is stored in your selected region. The supported region for this deployment scope is China (Beijing).
| Model name** | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota (Note) |
|---|---|---|---|
| deepseek-v4-pro ** Discounts are available for context caching. | $1.65 | $3.301 | No free quota |
| deepseek-v4-flash Discounts are available for context caching. | $0.138 | $0.275 | |
| deepseek-v3.2 Discounts are available for context caching. | $0.287 | $0.431 | |
| deepseek-v3.2-exp | $0.287 | $0.431 | |
| deepseek-v3.1 | $0.574 | $1.721 | |
| deepseek-r1 | $0.574 | $2.294 | |
| deepseek-r1-0528 | $0.574 | $2.294 | |
| deepseek-v3 | $0.287 | $1.147 | |
| deepseek-r1-distill-qwen-1.5b | Free for a limited time | ||
| deepseek-r1-distill-qwen-7b | $0.072 | $0.144 | No free quota |
| deepseek-r1-distill-qwen-14b | $0.144 | $0.431 | |
| deepseek-r1-distill-qwen-32b | $0.287 | $0.861 | |
| deepseek-r1-distill-llama-8b | Free for a limited time | ||
| deepseek-r1-distill-llama-70b | Free for a limited time |
Kimi
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model name** | Mode | Input price (per 1M tokens) | Output price (per 1M tokens) | Free quota(Note) |
|---|---|---|---|---|
| kimi-k2.6 | thinking and non-thinking modes | $0.8939 | $3.7131 | No free quota |
| kimi-k2.5 | thinking and non-thinking modes | $0.574 | $3.011 | |
| kimi-k2-thinking | thinking mode only | $0.574 | $2.294 | |
| Moonshot-Kimi-K2-Instruct | non-thinking mode | $0.574 | $2.294 |
MiniMax
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland. You are charged for input tokens and output tokens.
| Model name | Mode | Input price (per 1M tokens) | Output price (per 1M tokens) ** Chain of Thought (CoT) and response** |
|---|---|---|---|
| MiniMax-M2.5 | thinking mode only | $0.304 | $1.213 |
GLM
You are charged for input tokens and output tokens.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Model | Request input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** Chain of Thought (CoT) and response** |
|---|---|---|---|---|
| glm-5.1 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.825 | $3.301 |
| 32K < tokens ≤ 200K | $1.100 | $3.851 |
Chinese mainland
With the Chinese mainland deployment scope, model inference resources are restricted to the Chinese mainland, and static data is stored in your selected region. This scope is available in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Model | Request input tokens | Input price (per 1M tokens) | Output price (per 1M tokens) ** Chain of Thought (CoT) and response** |
|---|---|---|---|---|
| glm-5.1 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.825 | $3.301 |
| 32K < tokens ≤ 200K | $1.100 | $3.851 | ||
| glm-5 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.573 | $2.58 |
| 32K < tokens ≤ 166K | $0.86 | $3.154 | ||
| glm-4.7 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.431 | $2.007 |
| 32K < tokens ≤ 166K | $0.574 | $2.294 | ||
| glm-4.6 | thinking and non-thinking modes | 0 < tokens ≤ 32K | $0.431 | $2.007 |
| 32K < tokens ≤ 166K | $0.574 | $2.294 |
Image generation
You are not charged for input. You are charged for output based on the number of successfully generated images.
Formula: Cost = Image unit price × Number of images generated.
Notes:
Cost does not depend on image resolution or aspect ratio.
Failed requests incur no cost and do not consume your free quota.
Billing example: Some images fail to generate Assume the image unit price is $0.10 per image. If you call the API to generate four images but only three image URLs return successfully, the system charges only for the three successfully generated images.
Number billed: 3 images.
Cost calculation: 0.1 × 3 = $0.3.
Qwen text-to-image
You are charged for output only. For billing rules, see Image generation.
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model | Output price | Free quota (Note) |
|---|---|---|
| qwen-image-2.0-pro | $0.075/image | 100 images per model. The free quota is valid for 90 days after you activate Model Studio. |
| qwen-image-2.0-pro-2026-04-22 | $0.075/image | |
| qwen-image-2.0-pro-2026-03-03 | $0.075/image | |
| qwen-image-2.0 | $0.035/image | |
| qwen-image-2.0-2026-03-03 | $0.035/image | |
| qwen-image-max ** Currently equivalent to qwen-image-max-2025-12-30 | $0.075/image | |
| qwen-image-max-2025-12-30 | $0.075/image | |
| qwen-image-plus Currently equivalent to qwen-image | $0.03/image | |
| qwen-image-plus-2026-01-09 | $0.03/image | |
| qwen-image | $0.035/image |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Output price |
|---|---|
| qwen-image-2.0-pro | $0.071676/image |
| qwen-image-2.0-pro-2026-04-22 | $0.071676/image |
| qwen-image-2.0-pro-2026-03-03 | $0.071676/image |
| qwen-image-2.0 | $0.028671/image |
| qwen-image-2.0-2026-03-03 | $0.028671/image |
| qwen-image-max ** Currently equivalent to qwen-image-max-2025-12-30 | $0.071677/image |
| qwen-image-max-2025-12-30 | $0.071677/image |
| qwen-image-plus Currently equivalent to qwen-image | $0.028671/image |
| qwen-image-plus-2026-01-09 | $0.028671/image |
| qwen-image | $0.035/image |
Qwen-Image-Edit
You are charged for output only. For billing rules, see Image generation.
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model** | Output price | Free quota (Note) |
|---|---|---|
| qwen-image-2.0-pro | $0.075/image | 100 images per model. The free quota is valid for 90 days after you activate Model Studio. |
| qwen-image-2.0-pro-2026-04-22 | $0.075/image | |
| qwen-image-2.0-pro-2026-03-03 | $0.075/image | |
| qwen-image-2.0 | $0.035/image | |
| qwen-image-2.0-2026-03-03 | $0.035/image | |
| qwen-image-edit-max ** Currently equivalent to qwen-image-edit-max-2026-01-16 | $0.075/image | |
| qwen-image-edit-max-2026-01-16 | $0.075/image | |
| qwen-image-edit-plus Currently equivalent to qwen-image-edit-plus-2025-10-30 | $0.03/image | |
| qwen-image-edit-plus-2025-12-15 | $0.03/image | |
| qwen-image-edit-plus-2025-10-30 | $0.03/image | |
| qwen-image-edit | $0.045/image |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model** | Output price |
|---|---|
| qwen-image-2.0-pro | $0.071676/image |
| qwen-image-2.0-pro-2026-04-22 | $0.071676/image |
| qwen-image-2.0-pro-2026-03-03 | $0.071676/image |
| qwen-image-2.0 | $0.028671/image |
| qwen-image-2.0-2026-03-03 | $0.028671/image |
| qwen-image-edit-max ** Currently equivalent to qwen-image-edit-max-2026-01-16 | $0.071677/image |
| qwen-image-edit-max-2026-01-16 | $0.071677/image |
| qwen-image-edit-plus Currently equivalent to qwen-image-edit-plus-2025-10-30 | $0.028671/image |
| qwen-image-edit-plus-2025-12-15 | $0.028671/image |
| qwen-image-edit-plus-2025-10-30 | $0.028671/image |
| qwen-image-edit | $0.043/image |
Qwen-MT-Image
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
You are charged for output only. For billing rules, see Image generation.
| Model** | Output price | Free quota (Note) |
|---|---|---|
| qwen-mt-image | $0.000431/image | No free quota |
Qwen-Z-Image text-to-image
You are charged for output only. For billing rules, see Image generation.
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model | Output price | Free quota (Note) |
|---|---|---|
| z-image-turbo | Disabling prompt rewrite (prompt_extend=false): $0.015/imageEnabling prompt rewrite (prompt_extend=true): $0.03/image | 100 imagesValidity period: 90 days after Alibaba Cloud Model Studio is activated. |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output price |
|---|---|
| z-image-turbo | Prompt rewrite disabled (prompt_extend=false): $0.01434 per imageEnabling prompt rewrite (prompt_extend=true): $0.02868 per image |
Wan text-to-image
You are charged for output only. For billing rules, see Image generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model | Output price |
|---|---|
| wan2.6-t2i | $0.028671/image |
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model | Output price | Free quota (Note) |
|---|---|---|
| wan2.6-t2i | $0.03/image | 50 images |
| wan2.5-t2i-preview | $0.03/image | 50 images |
| wan2.2-t2i-plus | $0.05/image | 100 images |
| wan2.2-t2i-flash | $0.025/image | 100 images |
| wan2.1-t2i-plus | $0.05/image | 200 images |
| wan2.1-t2i-turbo | $0.025/image | 200 images |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output price |
|---|---|
| wan2.6-t2i | $0.028671/image |
| wan2.5-t2i-preview | $0.028671/image |
| wan2.2-t2i-plus | $0.020070/image |
| wan2.2-t2i-flash | $0.028671/image |
| wanx2.1-t2i-plus | $0.028671/image |
| wanx2.1-t2i-turbo | $0.020070/image |
| wanx2.0-t2i-turbo | $0.005735/image |
Wan image generation and editing
You are charged for output only. For billing rules, see Image generation.
Global (Virginia)
Note
Models deployed globally (Virginia) do not offer a free quota.
| Model | Output price |
|---|---|
| wan2.6-image | $0.028671/image |
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model | Output price | Free quota (Note) |
|---|---|---|
| wan2.7-image-pro | $0.075/image | 50 images. The free quota is valid for 90 days after you activate Model Studio. |
| wan2.7-image | $0.03/image | 50 images. The free quota is valid for 90 days after you activate Model Studio. |
| wan2.6-image | $0.03/image | 50 images. The free quota is valid for 90 days after you activate Model Studio. |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output price |
|---|---|
| wan2.7-image-pro | $0.068761/image |
| wan2.7-image | $0.028671/image |
| wan2.6-image | $0.028671/image |
Wan general image editing
You are charged for output only. For billing rules, see Image generation.
International
When you select the International deployment scope, the service dynamically schedules model inference compute resources globally (excluding the Chinese mainland). Static data is stored in your selected region, which for this scope is Singapore.
| Model service | Model | Output price | Free quota (Note) |
|---|---|---|---|
| General Image Editing 2.5 | wan2.5-i2i-preview | $0.03/image | 50 images. The free quota is valid for 90 days after you activate Model Studio. |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model service | Model | Output price |
|---|---|---|
| General Image Editing 2.5 | wan2.5-i2i-preview | $0.028671/image |
| General Image Editing 2.1 | wanx2.1-imageedit | $0.020070/image |
OutfitAnyone
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
aitryon-plus: You are charged for output only. For billing rules, see Image generation.
aitryon-parsing-v1: You are charged per input image, not per output image. Failed requests do not incur charges.
| Model service | Model | Price | Free quota (Note) |
|---|---|---|---|
| OutfitAnyone - Plus | aitryon-plus | $0.071677/image | No free quota |
| OutfitAnyone - Image Parsing | aitryon-parsing-v1 | $0.000574/image |
Video generation
You are not charged for input. You are charged for output based on the total duration of successfully generated videos (in seconds).
Formula: Cost = Video unit price × Video duration (seconds).
Notes:
Some models charge by output video resolution. Prices differ for resolutions such as 480P, 720P, and 1080P.
Some models charge by output video edition. Prices differ for editions such as Standard Edition and Professional Edition.
Some models charge by output video aspect ratio. Prices differ for aspect ratios such as 1:1 and 3:4.
Some models use a flat rate, regardless of resolution, edition, or aspect ratio.
Failed requests incur no cost and do not consume your free quota.
HappyHorse - Text-to-video
Billing is based on output only. For billing rules, see video generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model | Output resolution | Unit price |
|---|---|---|
| happyhorse-1.0-t2v | 720p | $0.123769/second |
| 1080p | $0.220034/second |
International
In the International deployment scope, model inference compute resources are dynamically scheduled globally, excluding the Chinese mainland. Static data is stored in your selected region. The supported region is Singapore.
| Model | Output resolution | Unit price | Free Tier (Note) Valid for 90 days after activating Model Studio |
|---|---|---|---|
| happyhorse-1.0-t2v | 720p | $0.14/second | 10 seconds |
| 1080p | $0.24/second |
Chinese mainland
In the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output resolution | Unit price |
|---|---|---|
| happyhorse-1.0-t2v | 720p | $0.123769/second |
| 1080p | $0.220034/second |
HappyHorse - image-to-video (first frame)
You are charged for output only. For billing rules, see video generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-i2v | 720P | $0.123769/second |
| 1080P | $0.220034/second |
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in Singapore, the only supported region for this scope.
| Model name | Output video resolution | Unit price | Free quota (Note) Valid for 90 days after you activate Model Studio. |
|---|---|---|---|
| happyhorse-1.0-i2v | 720P | $0.14/second | 10 seconds |
| 1080P | $0.24/second |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in China (Beijing), the only supported region for this scope. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-i2v | 720P | $0.123769/second |
| 1080P | $0.220034/second |
HappyHorse - Reference to video
You are billed only for the output. For billing rules, see video generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-r2v | 720p | $0.123769/second |
| 1080p | $0.220034/second |
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the region you select. Supported region: Singapore.
| Model name | Output video resolution | Unit price | Free quota(Note) Validity: 90 days after activating Model Studio |
|---|---|---|---|
| happyhorse-1.0-r2v | 720p | $0.14/second | 10 seconds |
| 1080p | $0.24/second |
Chinese mainland
When the deployment scope is Mainland China, model inference computing resources are limited to Mainland China, and static data is stored in the region that you select. The supported region for this deployment scope is: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-r2v | 720p | $0.123769/second |
| 1080p | $0.220034/second |
HappyHorse - video editing
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
Billing rules: You are billed for both input and output videos per second. Failed requests are not billed.
| Model | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-video-edit | 720p | $0.123769/second |
| 1080p | $0.220034/second |
International
The International deployment scope dynamically schedules model inference compute resources worldwide (excluding the Chinese mainland). Static data is stored in the only supported region: Singapore.
Billing rules: You are billed for both input and output videos per second. Failed requests are not billed and do not count against your free quota.
| Model | Output video resolution | Unit price | Free quota (Note) Validity: 90 days after you activate Model Studio |
|---|---|---|---|
| happyhorse-1.0-video-edit | 720p | $0.14/second | 10 seconds |
| 1080p | $0.24/second |
Chinese mainland
The Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in the only supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
Billing rules: You are billed for both input and output videos per second. Failed requests are not billed.
| Model | Output video resolution | Unit price |
|---|---|---|
| happyhorse-1.0-video-edit | 720p | $0.123769/second |
| 1080p | $0.220034/second |
Wan - text-to-video
You are billed for output only. For billing rules, see video generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| wan2.6-t2v | 720p | $0.086012/second |
| 1080p | $0.143353/second |
International
When you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland). Static data is stored in the Singapore (Singapore) region.
| Model name | Output video resolution | Unit price | Free quota (Note) Validity: 90 days after you activate Model Studio |
|---|---|---|---|
| wan2.7-t2v-2026-04-25 | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | ||
| wan2.7-t2v | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | ||
| wan2.6-t2v | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | ||
| wan2.5-t2v-preview | 480p | $0.05/second | 50 seconds |
| 720p | $0.10/second | ||
| 1080p | $0.15/second | ||
| wan2.2-t2v-plus | 480p | $0.02/second | 50 seconds |
| 1080p | $0.10/second | ||
| wan2.1-t2v-turbo | 480p | $0.036/second | 200 seconds |
| 720p | $0.036/second | ||
| wan2.1-t2v-plus | 720p | $0.10/second | 200 seconds |
US
When you select the US deployment scope, compute resources for model inference are restricted to the United States. Static data is stored in the US (Virginia) region. Note
Models in the US deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| wan2.6-t2v-us | 720p | $0.1/second |
| 1080p | $0.15/second |
Chinese mainland
When you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland. Static data is stored in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output video resolution | Unit price |
|---|---|---|
| wan2.7-t2v-2026-04-25 | 720p | $0.086012/second |
| 1080p | $0.143353/second | |
| wan2.7-t2v | 720p | $0.086012/second |
| 1080p | $0.143353/second | |
| wan2.6-t2v | 720p | $0.086012/second |
| 1080p | $0.143353/second | |
| wan2.5-t2v-preview | 480p | $0.043006/second |
| 720p | $0.086012/second | |
| 1080p | $0.143353/second | |
| wan2.2-t2v-plus | 480p | $0.02007/second |
| 1080p | $0.100347/second | |
| wan2.1-t2v-turbo | 480p | $0.034405/second |
| 720p | $0.034405/second | |
| wan2.1-t2v-plus | 720p | $0.100347/second |
Wan - image-to-video
You are billed for output only. For billing rules, see video generation.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is Singapore.
| Model | Output video type | Output video resolution | Price | Free quota (Note) Valid for 90 days after activating Model Studio |
|---|---|---|---|---|
| wan2.7-i2v-2026-04-25 | Video with audio | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | |||
| wan2.7-i2v | Video with audio | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output video type | Output video resolution | Price |
|---|---|---|---|
| wan2.7-i2v-2026-04-25 | Video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second | ||
| wan2.7-i2v | Video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second |
Wan: Image-to-video (from first frame)
You are billed for output only. For billing rules, see video generation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Output type | Output video resolution | Unit price |
|---|---|---|---|
| wan2.6-i2v | video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second |
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data resides in your selected region. Supported region: Singapore.
| Model name | Output type | Output video resolution | Unit price | Free quota (Note) Valid for 90 days after activating Model Studio |
|---|---|---|---|---|
| wan2.6-i2v-flash | video with audioaudio=true | 720p | $0.05/second | 50 seconds |
| 1080p | $0.075/second | |||
video without audioaudio=false | 720p | $0.025/second | ||
| 1080p | $0.0375/second | |||
| wan2.6-i2v | video with audio | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | |||
| wan2.5-i2v-preview | video with audio | 480p | $0.05/second | 50 seconds |
| 720p | $0.10/second | |||
| 1080p | $0.15/second | |||
| wan2.2-i2v-flash | video without audio | 480p | $0.015/second | 50 seconds |
| 720p | $0.036/second | |||
| wan2.2-i2v-plus | video without audio | 480p | $0.02/second | 50 seconds |
| 1080p | $0.10/second | |||
| wan2.1-t2v-turbo | video without audio | 480p | $0.036/second | 200 seconds |
| 720p | $0.036/second | |||
| wan2.1-t2v-plus | video without audio | 720p | $0.10/second | 200 seconds |
US
If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data resides in your selected region. Supported region: US (Virginia). Note
Models in the US deployment scope do not offer a free quota.
| Model name | Output type | Output video resolution | Unit price |
|---|---|---|---|
| wan2.6-i2v-us | video with audio | 720p | $0.1/second |
| 1080p | $0.15/second |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data resides in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output type | Output video resolution | Unit price |
|---|---|---|---|
| wan2.6-i2v-flash | video with audioaudio=true | 720p | $0.043006/second |
| 1080p | $0.071676/second | ||
video without audioaudio=false | 720p | $0.021503/second | |
| 1080p | $0.035838/second | ||
| wan2.6-i2v | video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second | ||
| wan2.5-i2v-preview | video with audio | 480p | $0.043006/second |
| 720p | $0.086012/second | ||
| 1080p | $0.143353/second | ||
| wan2.2-i2v-plus | video without audio | 480p | $0.02007/second |
| 1080p | $0.100347/second | ||
| wanx2.1-t2v-turbo | video without audio | 480p | $0.034405/second |
| 720p | $0.034405/second | ||
| wanx2.1-t2v-plus | video without audio | 720p | $0.100347/second |
Wan - Image-to-Video (Start and End Frames)
You are billed for output only. For billing rules, see video generation.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data resides in the region you select. The supported region for this deployment scope is Singapore.
| Model | Output resolution | Output price | Free quota (Note) Valid for 90 days after you activate Alibaba Cloud Model Studio |
|---|---|---|---|
| wan2.2-kf2v-flash | 480p | $0.015/second | 50 seconds |
| 720p | $0.036/second | ||
| 1080p | $0.07/second | ||
| wan2.1-kf2v-plus | 720p | $0.10/second | 200 seconds |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data resides in the region you select. The supported region for this deployment scope is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output resolution | Output price |
|---|---|---|
| wan2.2-kf2v-flash | 480p | $0.014335/second |
| 720p | $0.028671/second | |
| 1080p | $0.068809/second | |
| wanx2.1-kf2v-plus | 720p | $0.100347/second |
Wan - reference-to-video
Billing rules: You are billed for both input and output videos based on their duration in seconds. Failed requests do not incur charges or consume your free quota.
Billing formula: Billed duration = Input video duration (up to 5 seconds) + Output video duration.
The billed duration for the input video is capped at 5 seconds . For calculation rules, see Billing and rate limits.
The billed duration for the output video is the duration in seconds of the successfully generated video.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt). Note
Models in the global deployment scope do not offer a free quota.
| Model name | Output video type | Output video resolution | Input and output price |
|---|---|---|---|
| wan2.6-r2v | video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second |
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.
| Model name | Output video type | Output video resolution | Input and output price | Free quota (Note) Validity: 90 days after activating Model Studio |
|---|---|---|---|---|
| wan2.7-r2v | video with audio | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second | |||
| wan2.6-r2v-flash | video with audioaudio=true | 720p | $0.05/second | 50 seconds |
| 1080p | $0.075/second | |||
video without audioaudio=false | 720p | $0.025/second | ||
| 1080p | $0.0375/second | |||
| wan2.6-r2v | video with audio | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output video type | Output video resolution | Input and output price |
|---|---|---|---|
| wan2.7-r2v | video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second | ||
| wan2.6-r2v-flash | video with audioaudio=true | 720p | $0.043006/second |
| 1080p | $0.071676/second | ||
video without audioaudio=false | 720p | $0.021503/second | |
| 1080p | $0.035838/second | ||
| wan2.6-r2v | video with audio | 720p | $0.086012/second |
| 1080p | $0.143353/second |
Wan - video editing
International
When the deployment scope is International, model inference computing resources are dynamically scheduled globally (excluding mainland China). Static data is stored in the region you select. The supported region for this deployment scope is Singapore.
You are charged for both input and output videos per second of video duration. Failed tasks are not billed and do not consume your free quota.
| Model name | Output video resolution | Combined price | Free quota (Note) Valid for 90 days after you activate Model Studio. |
|---|---|---|---|
| wan2.7-videoedit | 720p | $0.10/second | 50 seconds |
| 1080p | $0.15/second |
You are charged for output videos per second of video duration. Failed tasks are not billed and do not consume your free quota.
| Model name | Output video resolution | Output price | Free quota (Note) Valid for 90 days after you activate Model Studio. |
|---|---|---|---|
| wan2.1-vace-plus | 720p | $0.10/second | 50 seconds |
Chinese mainland
When the deployment scope is mainland China, model inference computing resources are restricted to mainland China, and static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
You are charged for both input and output videos per second of video duration. You are not charged for failed tasks.
| Model name | Output video resolution | Combined price |
|---|---|---|
| wan2.7-videoedit | 720p | $0.086012/second |
| 1080p | $0.143353/second |
You are charged for output videos per second of video duration. You are not charged for failed tasks.
| Model name | Output video resolution | Output price |
|---|---|---|
| wanx2.1-vace-plus | 720p | $0.100347/second |
Wan - digital human
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
wan2.2-s2v-detect: You are charged for each input image per successful request, regardless of the detection result. There is no charge for output.wan2.2-s2v: You are billed per second for each successfully generated output video. There is no charge for input. For billing rules, see video generation.
| Service | Model name | Price | Free quota (Note) |
|---|---|---|---|
| Image detection | wan2.2-s2v-detect | Input image: $0.000574/image | No free quota |
| Video generation | wan2.2-s2v | Output video: - 480P: $0.071677/second - 720P: $0.129018/second |
Wan - image-to-motion
You are billed for output only. For billing rules, see video generation.
International
If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Singapore is the only supported region for this deployment scope.
| Model | Output video mode | Output price | Free quotaRules |
|---|---|---|---|
| wan2.2-animate-move | standard mode wan-std | $0.12/second | 50 secondsValid for 90 days after you activate Model Studio. |
professional mode wan-pro | $0.18/second |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. China (Beijing) is the only supported region for this deployment scope. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Output video mode | Output price |
|---|---|---|
| wan2.2-animate-move | standard mode wan-std | $0.06/second |
professional mode wan-pro | $0.09/second |
Wan - video character swap
You are billed for output only. For billing rules, see video generation.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in the selected region, which must be Singapore.
| Model name | Output mode | Price | Free Tier (Note) |
|---|---|---|---|
| wan2.2-animate-mix | Standard mode wan-std | $0.18/second | 50 secondsThe free quota is valid for 90 days after you activate Model Studio. |
Professional modewan-pro | $0.26/second |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in the selected region, which must be China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Output mode | Price |
|---|---|---|
| wan2.2-animate-mix | Standard mode wan-std | $0.09/second |
Professional modewan-pro | $0.13/second |
AnimateAnyone
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
animate-anyone-detect-gen2: Charges are based on the input image. You are billed once for each input image per successful request, regardless of the detection outcome.animate-anyone-template-gen2: Charges are based on the output video. You are billed per second for each successfully generated video. For billing rules, see video generation.animate-anyone-gen2: Charges are based on the output video. You are billed per second for each successfully generated video. For billing rules, see video generation.
| Model service | Model | Price | Free quota(Note) |
|---|---|---|---|
| Image detection | animate-anyone-detect-gen2 | $0.000574 per input image | No free quota |
| Action template generation | animate-anyone-template-gen2 | Output video: $0.011469/second | |
| Video generation | animate-anyone-gen2 | Output video: $0.011469/second |
EMO
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
emo-detect-v1: Billing is based on input, not output. You are charged for each input image per successful request, regardless of the detection result.emo-v1: Billing is based on output, not input. You are charged per second of successfully generated video. For billing rules, see video generation.
| Model service | Model name | Price | Free quota(note) |
|---|---|---|---|
| image detection | emo-detect-v1 | input image: $0.000574/image | No free quota |
| video generation | emo-v1 | output video: - Video (1:1 aspect ratio): $0.011469/second - Video (3:4 aspect ratio): $0.022937/second |
LivePortrait
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
liveportrait-detect: You are charged for input but not for output. Each successful request is charged once per input image, regardless of the detection result.liveportrait: You are charged for output but not for input. You are billed per second for each successfully generated video. For more information about billing rules, see video generation.
| Service | Model | Price | Free quota (Note) |
|---|---|---|---|
| Image detection | liveportrait-detect | Input image: $0.000574 per image | No free quota is provided. |
| Video generation | liveportrait | Output video: $0.002868 per second |
Emoji
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
emoji-detect-v1: You are charged for the input, not the output. This charge applies once per input image for each successful request, regardless of the detection result.emoji-v1: You are not charged for the input. You are billed per second of successfully generated video. For billing rules, see video generation.
| Model service | Model name | Unit price | Free quota (Note) |
|---|---|---|---|
| Image detection | emoji-detect-v1 | $0.000574 per input image | No free quota |
| Video generation | emoji-v1 | $0.011469 per second |
VideoRetalk
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
You are billed for output only. For billing rules, see video generation.
| Model name | Output price | Free quota (Note) |
|---|---|---|
| videoretalk | $0.011469/second | No free quota |
Video style transform
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
You are billed for output only. For billing rules, see video generation.
| Model name | Output resolution | Unit price | Free quota Note |
|---|---|---|---|
| video-style-transform | 540p | $0.028671/second | No free quota |
| 720p | $0.071677/second |
Music generation
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
Billing rule: You are billed only for the duration (in seconds) of the output audio.
| Model | Unit price (per second) | Free quota(Note) |
|---|---|---|
| fun-music-v1 | $0.000275 | No free quota |
Speech synthesis (text-to-speech)
Qwen-TTS
International
When the deployment scope is International, model inference compute resources are dynamically scheduled globally (excluding mainland China). Static data is stored in the region that you select. The supported region for this deployment scope is Singapore.
Qwen3-TTS-Instruct-Flash
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| qwen3-tts-instruct-flash | $0.115 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-instruct-flash-2026-01-26 | $0.115 |
Qwen3-TTS-VD
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| qwen3-tts-vd-2026-01-26 | $0.115 | 10,000 charactersValidity: 90 days after activating Model Studio |
Qwen3-TTS-VC
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| qwen3-tts-vc-2026-01-22 | $0.115 | 10,000 charactersValidity: 90 days after activating Model Studio |
Qwen3-TTS-Flash
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| qwen3-tts-flash ** Currently equivalent to qwen3-tts-flash-2025-11-27 | $0.1 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-flash-2025-11-27 | $0.1 | |
| qwen3-tts-flash-2025-09-18 | $0.1 | For activations before 00:00, November 13, 2025: 2,000 charactersFor activations at or after 00:00, November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio |
Chinese mainland
When the deployment scope is mainland China, model inference computing resources are limited to mainland China, and static data is stored in the region that you select. The region supported by this deployment scope is: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
Qwen3-TTS-Instruct-Flash
You are billed based on the number of characters in the input text; the output is not charged.
| Model name** | Input price (per 10,000 characters) | Output price (per 10,000 characters) |
|---|---|---|
| qwen3-tts-instruct-flash | $0.115 | Not billed |
| qwen3-tts-instruct-flash-2026-01-26 | $0.115 | Not billed |
Qwen3-TTS-VD
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Output price (per 10,000 characters) |
|---|---|---|
| qwen3-tts-vd-2026-01-26 | $0.115 | Not billed |
Qwen3-TTS-VC
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Output price (per 10,000 characters) |
|---|---|---|
| qwen3-tts-vc-2026-01-22 | $0.115 | Not billed |
Qwen3-TTS-Flash
You are billed based on the number of characters in the input text; the output is not charged.
| Model name | Input price (per 10,000 characters) | Output price (per 10,000 characters) |
|---|---|---|
| qwen3-tts-flash ** Currently equivalent to qwen3-tts-flash-2025-11-27 | $0.114682 | Not billed |
| qwen3-tts-flash-2025-11-27 | $0.114682 | Not billed |
| qwen3-tts-flash-2025-09-18 | $0.114682 | Not billed |
Qwen-TTS
You are billed based on the number of input and output tokens.
| Model name** | Input price (per million tokens) | Output price (per million tokens) |
|---|---|---|
| qwen-tts-flash | $0.23 | $1.434 |
| qwen-tts-latest | $0.23 | $1.434 |
| qwen-tts-2025-05-22 | $0.23 | $1.434 |
| qwen-tts-2025-04-10 | $0.23 | $1.434 |
Qwen-TTS-Realtime
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in the region you select. The supported region is Singapore.
Qwen3-TTS-Instruct-Flash-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Free quota (note) |
|---|---|---|
| qwen3-tts-instruct-flash-realtime | $0.143 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-instruct-flash-realtime-2026-01-22 | $0.143 | 10,000 charactersValidity: 90 days after activating Model Studio |
Qwen3-TTS-VD-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Free quota (note) |
|---|---|---|
| qwen3-tts-vd-realtime-2026-01-15 | $0.143353 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-vd-realtime-2025-12-16 | $0.143353 | 10,000 charactersValidity: 90 days after activating Model Studio |
Qwen3-TTS-VC-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Free quota (note) |
|---|---|---|
| qwen3-tts-vc-realtime-2026-01-15 | $0.13 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-vc-realtime-2025-11-27 |
Qwen3-TTS-Flash-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Free quota (note) |
|---|---|---|
| qwen3-tts-flash-realtime | $0.13 | For Model Studio activated before 00:00 on November 13, 2025: 2,000 charactersFor Model Studio activated at or after 00:00 on November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-flash-realtime-2025-11-27 | $0.13 | 10,000 charactersValidity: 90 days after activating Model Studio |
| qwen3-tts-flash-realtime-2025-09-18 | $0.13 | For Model Studio activated before 00:00 on November 13, 2025: 2,000 charactersFor Model Studio activated at or after 00:00 on November 13, 2025: 10,000 charactersValidity: 90 days after activating Model Studio |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in the region you select. The supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
Qwen3-TTS-Instruct-Flash-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Output price |
|---|---|---|
| qwen3-tts-instruct-flash-realtime | $0.143 | Not charged |
| qwen3-tts-instruct-flash-realtime-2026-01-22 | $0.143 | Not charged |
Qwen3-TTS-VD-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Output price |
|---|---|---|
| qwen3-tts-vd-realtime-2026-01-15 | $0.143353 | Not charged |
| qwen3-tts-vd-realtime-2025-12-16 | $0.143353 | Not charged |
Qwen3-TTS-VC-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Output price |
|---|---|---|
| qwen3-tts-vc-realtime-2026-01-15 | $0.143353 | Not charged |
| qwen3-tts-vc-realtime-2025-11-27 |
Qwen3-TTS-Flash-Realtime
Billing rules: Billing is based on the number of characters in the input text. Output is not billed.
| Model | Input price / 10,000 characters | Output price |
|---|---|---|
| qwen3-tts-flash-realtime | $0.143353 | Not charged |
| qwen3-tts-flash-realtime-2025-11-27 | $0.143353 | Not charged |
| qwen3-tts-flash-realtime-2025-09-18 | $0.143353 | Not charged |
Qwen-TTS-Realtime
Billing rules: Billing is based on the number of input and output tokens.
| Model | Input price / 1,000,000 tokens | Output price / 1,000,000 tokens |
|---|---|---|
| qwen-tts-realtime | $0.345 | $1.721 |
| qwen-tts-realtime-latest | $0.345 | $1.721 |
| qwen-tts-realtime-2025-07-15 | $0.345 | $1.721 |
Qwen-TTS voice cloning
Billing: You are charged for each new voice.
International
The International deployment scope dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in Singapore.
| Model name | Price (per voice) | Free quota (Note) |
|---|---|---|
| qwen-voice-enrollment | $0.01 | 1,000 voices/account |
Chinese mainland
The Chinese mainland deployment scope restricts model inference compute resources to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Price (per voice) |
|---|---|
| qwen-voice-enrollment | $0.01 |
Qwen-TTS voice design
You are charged for each new voice you create.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in Singapore.
| Model | Price (per voice) | Free quota(Note) |
|---|---|---|
| qwen-voice-design | $0.2 | 10 voices/account |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. This deployment scope is available in China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Price (per voice) |
|---|---|
| qwen-voice-design | $0.2 |
CosyVoice
International
If you select the International deployment scope, compute resources for model inference are dynamically scheduled worldwide (excluding the Chinese mainland), and your static data is stored in your selected region. The only supported region is Singapore.
Billing is based on the number of input characters; output is not charged.
| Model | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| cosyvoice-v3-plus | $0.26 | 10,000 charactersThe quota is valid for 90 days after you activate Model Studio. |
| cosyvoice-v3-flash | $0.13 |
Chinese mainland
If you select the Chinese mainland deployment scope, compute resources for model inference are restricted to the Chinese mainland, and your static data is stored in your selected region. The only supported region is China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
Billing is based on the number of input characters; output is not charged.
| Model | Input price (per 10,000 characters) | Free quota (Note) |
|---|---|---|
| cosyvoice-v3.5-plus | $0.22 | No free quota |
| cosyvoice-v3.5-flash | $0.116 | |
| cosyvoice-v3-plus | $0.2867 | |
| cosyvoice-v3-flash | $0.1434 | |
| cosyvoice-v2 | $0.2867 |
Speech recognition and speech translation
Qwen-LiveTranslate-Flash-Realtime
Billing rules: Billing is based on the number of input and output tokens. For details on how tokens are calculated for different modalities, see Billing.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.
| Model | Input price (per 1,000,000 tokens) | Output price (per 1,000,000 tokens) | Free quota (Note) | ||
|---|---|---|---|---|---|
| Input: Audio | Input: Image | Output: Text | Output: Audio | ||
| qwen3-livetranslate-flash-realtime | $10 | $1.3 | $10 | $38 | 1 million tokens for each modality Validity: 90 days after you activate Model Studio |
| qwen3-livetranslate-flash-realtime-2025-09-22 | $10 | $1.3 | $10 | $38 |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price (per 1,000,000 tokens) | Output price (per 1,000,000 tokens) | ||
|---|---|---|---|---|
| Input: Audio | Input: Image | Output: Text | Output: Audio | |
| qwen3-livetranslate-flash-realtime | $9.175 | $1.147 | $9.175 | $34.405 |
| qwen3-livetranslate-flash-realtime-2025-09-22 | $9.175 | $1.147 | $9.175 | $34.405 |
Qwen-ASR
Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.
| Model | Input price | Free quota (Note) |
|---|---|---|
| qwen3-asr-flash-filetrans | $0.000035/second | 36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio |
| qwen3-asr-flash-filetrans-2025-11-17 | ||
| qwen3-asr-flash ** Currently equivalent to qwen3-asr-flash-2025-09-08 | ||
| qwen3-asr-flash-2026-02-10 | ||
| qwen3-asr-flash-2025-09-08 |
US
When you select the US deployment scope, model inference compute resources are restricted to the United States. Your static data is stored in your selected region. This scope is available in the following region: US (Virginia). Note
Models in the US deployment scope do not offer a free quota.
| Model** | Input price |
|---|---|
| qwen3-asr-flash-us | $0.000035/second |
| qwen3-asr-flash-2025-09-08-us |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price |
|---|---|
| qwen3-asr-flash-filetrans | $0.000032/second |
| qwen3-asr-flash-filetrans-2025-11-17 | |
| qwen3-asr-flash ** Currently equivalent to qwen3-asr-flash-2025-09-08 | |
| qwen3-asr-flash-2026-02-10 | |
| qwen3-asr-flash-2025-09-08 |
Qwen-ASR-Realtime
Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.
| Model** | Input price | Free quota (Note) |
|---|---|---|
| qwen3-asr-flash-realtime | $0.000090/second | 36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio |
| qwen3-asr-flash-realtime-2026-02-10 | $0.000090/second | |
| qwen3-asr-flash-realtime-2025-10-27 | $0.000090/second |
Chinese mainland
When you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in your selected region. This scope is available in the following region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model | Input price |
|---|---|
| qwen3-asr-flash-realtime | $0.000047/second |
| qwen3-asr-flash-realtime-2026-02-10 | |
| qwen3-asr-flash-realtime-2025-10-27 |
Fun-ASR
Audio file recognition
Billing rules: Billing is based on the duration of the input audio in seconds. The output is not billed.
International
When you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide (excluding the Chinese mainland). Your static data is stored in your selected region. This scope is available in the following region: Singapore.
| Model | Input price | Free quota** (Note)** |
|---|---|---|
| fun-asr | $0.000035/second | 36,000 seconds (10 hours) Validity: 90 days after you activate Model Studio |
| fun-asr-2025-11-07 | ||
| fun-asr-2025-08-25 | ||
| fun-asr-mtl | ||
| fun-asr-mtl-2025-08-25 |
Mainland China
If the deployment scope is Mainland China, model inference compute resources are available only in Mainland China. Static data is stored in the region that you select. This deployment scope is supported in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model Name | Input Price |
|---|---|
| fun-asr | $0.000032/second |
| fun-asr-2025-11-07 | |
| fun-asr-2025-08-25 | |
| fun-asr-mtl | |
| fun-asr-mtl-2025-08-25 |
Real-time speech recognition
Billing rule: You are billed based on the duration of the input audio in seconds. The output is not billed.
International
When the deployment scope is set to International, model inference compute resources are dynamically scheduled worldwide, excluding mainland China. Static data is stored in your selected region. This deployment scope is supported in the Singapore region.
| Model name | Input price | Free tier** (Note)** |
|---|---|---|
| fun-asr-realtime | $0.00009/second | 36,000 seconds (10 hours)Valid for 90 days |
| fun-asr-realtime-2025-11-07 |
Mainland China
When the deployment scope is set to Mainland China, model inference compute resources are restricted to mainland China. Static data is stored in your selected region. This deployment scope is supported in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input price |
|---|---|
| fun-asr-realtime | $0.000047/second |
| fun-asr-realtime-2026-02-28 | |
| fun-asr-realtime-2025-11-07 | |
| fun-asr-realtime-2025-09-15 | |
| fun-asr-flash-8k-realtime | $0.000032/second |
| fun-asr-flash-8k-realtime-2026-01-28 |
Paraformer
Audio file recognition
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
Billing rules: You are billed based on the duration of the input audio in seconds. The output is not billed.
| Model | Input price |
|---|---|
| paraformer-v2 | $0.000012/second |
| paraformer-8k-v2 |
Real-time speech recognition
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
Billing rules: You are billed based on the duration of the input audio in seconds. The output is not billed.
| Model | Input price | Free tier(Note) |
|---|---|---|
| paraformer-realtime-v2 | $0.000035/second | No free tier |
| paraformer-realtime-8k-v2 |
Text embedding
Billing is based on input tokens. Output is not charged.
International
If you select the International deployment scope, model inference runs on dynamically scheduled resources worldwide, excluding the chinese mainland. Static data is stored in the region you select. Supported region: Singapore.
| Model name | Input price (per 1M tokens) | Free quota (Note) |
|---|---|---|
| text-embedding-v4 | $0.07 | 1,000,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio |
| text-embedding-v3 | $0.07 | 500,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio |
Chinese mainland
If you select the chinese mainland deployment scope, model inference is restricted to the chinese mainland. Static data is stored in the region you select. Supported region: China (Beijing). Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input price (per 1M tokens) |
|---|---|
| text-embedding-v4 | $0.072 |
China (Hong Kong)
If you select the China (Hong Kong) deployment scope, model inference is restricted to China (Hong Kong). Static data is stored in the region you select. Supported region: China (Hong Kong). Note
Models in the China (Hong Kong) deployment scope do not offer a free quota.
| Model name | Input price (per 1M tokens) | Free quota (Note) |
|---|---|---|
| text-embedding-v4 | $0.07 | 1,000,000 tokens Valid for 90 days after you activate Alibaba Cloud Model Studio |
Multimodal embedding
Billing is based on input tokens. Output is not charged.
International
If you select the International deployment scope, the system dynamically schedules model inference compute resources worldwide, excluding the Chinese mainland. The system stores static data in the region you select. Supported region: Singapore.
| Model | Input price | Free quota (Note) |
|---|---|---|
| tongyi-embedding-vision-plus | $0.09 | 1 million tokensValidity: 90 days after activating Model Studio |
| tongyi-embedding-vision-flash | Image/Video: $0.03Text: $0.09 |
Chinese mainland
If you select the Chinese mainland deployment scope, the system restricts model inference compute resources to the Chinese mainland. The system stores static data in the region you select. Supported region: China (Beijing).
| Model | Input price | Free quota (Note) |
|---|---|---|
| qwen3-vl-embedding | Image/Video: $0.258Text: $0.1 | 1 million tokensValidity: 90 days after activating Model Studio |
| multimodal-embedding-v1 | Free trial | Unlimited tokens |
Text rerank
Billing rule: You are charged for input tokens only.
International
For the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Your static data is stored in the Singapore region.
| Model name | Input price | Free quota (Note) |
|---|---|---|
| qwen3-rerank | $0.1 | 1 million tokensValid for 90 days after you activate Model Studio. |
Chinese mainland
For the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Your static data is stored in the China (Beijing) region. Note
Models in the Chinese mainland deployment scope do not offer a free quota.
| Model name | Input price |
|---|---|
| gte-rerank-v2 | $0.115 |
Domain-specific models
Intent recognition
Note
Only the Chinese mainland service deployment scope is supported. Data storage is in the Beijing access region. Model inference compute resources are limited to the Chinese mainland.
| Model name | Input price | Output price | Free quota(Note) |
|---|---|---|---|
| tongyi-intent-detect-v3 | $0.058 | $0.144 | No free quota |
Role playing
You are charged for input tokens and output tokens.
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. The service stores static data in the region you select. Supported region: Singapore.
| Model name | Input price | Output price | Free quota(Note) |
|---|---|---|---|
| qwen-plus-character ** Discounts apply for session cache. | $0.5 | $1.4 | 1 million tokens eachValid for 90 days after you activate Model Studio |
| qwen-flash-character Discounts apply for session cache. | $0.05 | $0.4 | |
| qwen-plus-character-ja | $0.5 | $1.4 |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. The service stores static data in the region you select. Supported region: China (Beijing).
| Model name** | Input price | Output price | Free quota(Note) |
|---|---|---|---|
| qwen-plus-character Discounts apply for session cache. | $0.115 | $0.287 | No free quota |
| qwen-flash-character Discounts apply for session cache. | $0.034 | $0.203 | No free quota |