Appearance
Model deployment
Deploy a to get a dedicated inference service that supports high concurrency and low latency.
Billing
Before deployment, view the estimated hourly cost for different models in the deployment console. Note
You cannot change the billing method after creating a service. To switch methods, undeploy and redeploy the model.
| Provisioned throughput (High throughput; high performance) | Model Unit (Custom performance metrics; resource isolation) | Token usage (Pay-as-you-go for fine-tuned models/performance validation) | |
|---|---|---|---|
| Definition | Reserves platform resources to guarantee a specific tokens per minute (TPM) throughput. No rate limiting within the guaranteed quota. | Provisions dedicated compute resources based on usage duration and number of model units. | Bills based on the number of input and output tokens per model call. |
| Advantages | - Provides stable throughput, lower latency, and predictable resource availability for high-load production environments. - Delivers 1.5 to 2.0 times the tokens per second (TPS) of the token usage method. - Supports auto-renewal. | - Customizable performance metrics such as latency and throughput. - Supports auto-renewal. | Pay only for what you use. |
| Supported models | Some preset models | Some preset models and all fine-tuned models | Some models fine-tuned with LoRA |
| Use cases | - Intelligent chatbots for banking apps with stable traffic and high concurrency. - Real-time content moderation for social media platforms handling predictable pipeline tasks. - Public cloud translation APIs providing baseline service guarantees for standard plan users. | - Custom fine-tuned large models for e-commerce, with manual scaling during high-traffic sales events. - Molecular screening models for pharmaceutical companies requiring dedicated resources for long-running tasks. - Autonomous driving simulations requiring continuous, long-term computation. | Validating the performance of fine-tuned models |
| Billing diagram | * | ||
| Billing method | Based on usage duration and provisioned throughput Pay-as-you-go or prepaid daily subscription | Based on usage duration and number of model units Pay-as-you-go or prepaid monthly subscription | Based on token usage Pay-as-you-go |
| Scaling method | Manually scale the provisioned throughput | Manually scale the number of model units | Submit a request for manual review in the console |
| Limitations | - Daily subscriptions are prepaid and non-refundable. - If your usage exceeds the provisioned throughput, the system automatically routes requests to the Model Studio model call service. | If you cancel a prepaid subscription within the first month, days used are billed at 1.2 times the standard daily rate. | - Supports only specific models fine-tuned with LoRA. - The system automatically releases resources after one month of inactivity. |
To view historical token usage and call counts, go to model monitoring.
Billing
Time-based billing (provisioned throughput)
Cost = Usage Duration × (Input TPM Unit Price × Input TPM + Output TPM Unit Price × Output TPM)
A subscription is activated upon payment and valid for N days, expiring at 23:59 on the Nth day. For orders placed after 22:00, the expiration date is extended by one day.
After a subscription expires, the service stops after a 2-hour grace period. Resources are retained for 14 hours before release.
You cannot terminate a subscription early.
For pay-as-you-go accounts with an overdue balance, deployed resources are retained and billed for 24 hours before automatic release.
If a model's input exceeds the max input tokens or the purchased TPM, the corresponding calls automatically switch to pay-as-you-go mode. In this case, inference performance may decrease, rate limiting reverts to the public traffic limits of the current snapshot model in your workspace, and fees are charged at the standard pay-as-you-go rate for model calls.
In this case, the API response header includes
x-dashscope-ptu-overflow:true.To view TPM statistics, go to model monitoring (China (Beijing)).
Singapore
| Model name | Model code | Thinking mode | Max input tokens | Pay-as-you-go (Hourly) | Subscription (Daily) | ||
|---|---|---|---|---|---|---|---|
| Input (per 10k TPM) | Output (per 1k TPM) | Input (per 10k TPM) | Output (per 1k TPM) | ||||
| DeepSeek-v3.2 | deepseek-v3.2 | Supported | 64,000 | $2.052 | $0.616 | $24.624 | $7.387 |
| Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | Supported | 128,000 | $1.200 | $0.720 | $14.400 | $8.640 |
| Qwen3.5-plus-2026-04-20 | qwen3.5-plus-2026-04-20 | Supported | 128,000 | $0.960 | $0.576 | $11.520 | $6.912 |
| Qwen3-vl-plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | Supported | 128,000 | $0.480 | $0.384 | $5.760 | $4.608 |
China (Beijing)
| Model name | Model code | Thinking mode | Max input tokens | Pay-as-you-go (Hourly) | Subscription (Daily) | ||
|---|---|---|---|---|---|---|---|
| Input (per 10k TPM) | Output (per 1k TPM) | Input (per 10k TPM) | Output (per 1k TPM) | ||||
| GLM-5.1 | glm-5.1 | Supported | 64,000 | $2.97 | $1.19 | $35.65 | $14.26 |
| Qwen3.6-Plus-2026-04-02 | qwen3.6-plus-2026-04-02 | Supported | 128,000 | $0.67 | $0.397 | $7.93 | $4.753 |
| Qwen3.5-plus-2026-04-20 | qwen3.5-plus-2026-04-20 | Supported | 128,000 | $0.26 | $0.16 | $3.17 | $1.9 |
| Qwen3-max-2025-09-23 | qwen3-max-2025-09-23 | Not supported | 128,000 | $1.11 | $0.45 | $13.32 | $5.40 |
| Qwen-plus-2025-12-01 | qwen-plus-2025-12-01 | Not supported | $0.28 | $0.07 | $3.36 | $0.84 | |
| Supported | $0.28 | $3.36 | |||||
| Qwen-flash-2025-07-28 | qwen-flash-2025-07-28 | Supported | $0.06 | $0.06 | $0.72 | $0.72 | |
| Qwen3-vl-plus-2025-09-23 | qwen3-vl-plus-2025-09-23 | Supported | $0.35 | $0.35 | $4.20 | $4.20 | |
| DeepSeek-v3.2 | deepseek-v3.2 | Supported | 64,000 | $1.04 | $0.16 | $12.48 | $1.92 |
Pay-as-you-go (model unit)
Fee = Usage (hours) × Model units × Price per model unit
- If you cancel a prepaid purchase within the first month, you are billed at 1.2 times the daily rate (any partial day is billed as a full day).
Note
Computing resources for pay-as-you-go model units are allocated on a first-come, first-served basis. If your purchase is unsuccessful, you receive a full refund.
Singapore
| Model name | Model code | Type | Rate limiting | Model unit | Max context | Unit price (Billed per minute; partial minutes are rounded up.) | Monthly price (Billed per day; partial days are rounded up.) (For subscriptions canceled within the first month, used days are billed at 1.2 times the standard daily rate.) |
|---|---|---|---|---|---|---|---|
| Qwen3-32B | qwen3-32b | Instruct | Yes | Type I model unit (MU1) | Fixed: 131,072 | $44/hour | $20,916/month |
Model type:
- Instruct: Performs inference in non-thinking mode.
China (Beijing)
| Model name | Model code | Type | Rate limiting | Model unit | Max context | Unit price (Billed per minute; partial minutes are rounded up.) | Monthly price (Billed per day; partial days are rounded up.) (For subscriptions canceled within the first month, used days are billed at 1.2 times the standard daily rate.) |
|---|---|---|---|---|---|---|---|
| Qwen3-14B | qwen3-14b | Instruct | Yes | Type I model unit (MU1) | Fixed: 131,072 | $40/hour | $18,800/month |
| Qwen3-32B | qwen3-32b | Instruct | Yes | Type I model unit (MU1) | Fixed: 131,072 | $40/hour | $18,800/month |
| Qwen3-VL-8B-Instruct | qwen3-vl-8b-instruct | Instruct | Yes | Type I model unit (MU1) | Fixed: 131,072 | $20/hour | $9,400/month |
| Qwen3-VL-8B-Thinking | qwen3-vl-8b-thinking | Thinking | Yes |
Model type:
Instruct: Performs inference in non-thinking mode.
Thinking: Performs inference in thinking mode.
Token-based billing
Cost = (Number of input tokens × Input unit price) + (Number of output tokens × Output unit price) (Minimum billing unit: 1 token)
- Token-based billing applies only to custom models fine-tuned from the following foundation models using Supervised Fine-Tuning (SFT).
Singapore
| Foundation model | Model code | Model type | Max context | Input unit price | Output unit price |
|---|---|---|---|---|---|
| Qwen3-14B | qwen3-14b-instruct | Instruct | Fixed at 131,072 | ¥0.001/1,000 tokens | Non-thinking mode: ¥0.004/1,000 tokens Thinking mode: ¥0.01/1,000 tokens |
To deploy additional models, consult this solution to select the best deployment plan for your business.
Deployment method
To deploy a model in the console, follow these steps:
If you receive an "insufficient permissions" error, see Deployment permission errors.
| - Go to the model deployment console (China (Beijing)). | |
|---|---|
| - Select a model and billing method, leave other settings as default, specify a model name, and start deployment. | |
| - A Running status indicates a successful deployment. * **Important ** Billing starts after the model deploys successfully. |
Deployment configuration
Model unit
| Parameter | Description |
|---|---|
| model inference mode | When deploying certain models as a Model Unit, you can configure the inference mode and max context. - Instruct - Runs in non-thinking mode. - Thinking - Runs in thinking mode. |
| max context | This setting is available for certain models deployed as a Model Unit. Max context length depends on the model type. |
| service throttling | This setting is available for certain models deployed as a Model Unit. You can limit the requests per minute (RPM) and tokens per minute (TPM) for model calls. |
Call a deployed model
After deploying a model, call it using the OpenAI-compatible API, DashScope, or the Assistant SDK.
When calling a deployed model, set the model parameter to the model code. Find the Model Code in the deployment console (China (Beijing)).
DashScope
HELPCODEESCAPE-python
import os
import dashscope
messages = [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who are you?"},
]
dashscope.base_http_api_url = 'https://dashscope-intl.aliyuncs.com/api/v1'
response = dashscope.Generation.call(
# If the DASHSCOPE_API_KEY environment variable is not set, provide your Model Studio API key directly, for example: api_key="sk-xxx"
api_key=os.getenv("DASHSCOPE_API_KEY"),
model="qwen3-max-xxx-xxx", # Replace with the model code of your deployed model.
messages=messages,
result_format="message",
enable_thinking=False,
)
print(response)OpenAI-compatible API
HELPCODEESCAPE-python
import os
from openai import OpenAI
client = OpenAI(
# If the DASHSCOPE_API_KEY environment variable is not set, provide your Model Studio API key directly, for example: api_key="sk-xxx"
api_key=os.getenv('DASHSCOPE_API_KEY'),
base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
)
completion = client.chat.completions.create(
model="qwen3-max-xxx-xxx", # Replace with the model code of your deployed model.
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Who are you?"},
],
extra_body={"enable_thinking": False},
)
print(completion)Service scaling
Provisioned throughput (time-based billing): Click Scaling to adjust the number of instances.
Model unit (time-based billing): Click Scaling to adjust the number of instances.
Service deactivation
Go to the model deployment console (China (Beijing)), find the deployment service to deactivate, and click Deactivate. After confirmation, billing for the service stops.
FAQ
Deploying your own models
Currently, you cannot upload and deploy your own models. Check Model Studio for the latest updates.
Alternatively, use Alibaba Cloud Platform for AI (PAI) to deploy your own models. For instructions, see Deploy large language models in PAI.
Deployment permission errors
If you see the error message "You Do Not Have Permissions For This Module" , ensure your account has the ModelDeploy-FullAccess permission for the workspace.
If you cannot proceed, contact your organization or IT administrator to grant the necessary permissions or investigate the issue.
If you receive the error message "xx workspace does not have deployment privilege for model xx " during deployment, go to the Model Studio Workspaces page and authorize the workspace to deploy the model.
API call error:
Workspace xxx does not have deployment privilege for model xxxx.If you still receive permission errors, contact your organization or IT administrator to grant the necessary permissions or perform the operation on your behalf.