Skip to content

Model deployment

Deploy a to get a dedicated inference service that supports high concurrency and low latency.

Billing

Before deployment, view the estimated hourly cost for different models in the deployment console. Note

You cannot change the billing method after creating a service. To switch methods, undeploy and redeploy the model.

Provisioned throughput (High throughput; high performance)Model Unit (Custom performance metrics; resource isolation)Token usage (Pay-as-you-go for fine-tuned models/performance validation)
DefinitionReserves platform resources to guarantee a specific tokens per minute (TPM) throughput. No rate limiting within the guaranteed quota.Provisions dedicated compute resources based on usage duration and number of model units.Bills based on the number of input and output tokens per model call.
Advantages- Provides stable throughput, lower latency, and predictable resource availability for high-load production environments. - Delivers 1.5 to 2.0 times the tokens per second (TPS) of the token usage method. - Supports auto-renewal.- Customizable performance metrics such as latency and throughput. - Supports auto-renewal.Pay only for what you use.
Supported modelsSome preset modelsSome preset models and all fine-tuned modelsSome models fine-tuned with LoRA
Use cases- Intelligent chatbots for banking apps with stable traffic and high concurrency. - Real-time content moderation for social media platforms handling predictable pipeline tasks. - Public cloud translation APIs providing baseline service guarantees for standard plan users.- Custom fine-tuned large models for e-commerce, with manual scaling during high-traffic sales events. - Molecular screening models for pharmaceutical companies requiring dedicated resources for long-running tasks. - Autonomous driving simulations requiring continuous, long-term computation.Validating the performance of fine-tuned models
Billing diagram*
Billing methodBased on usage duration and provisioned throughput Pay-as-you-go or prepaid daily subscriptionBased on usage duration and number of model units Pay-as-you-go or prepaid monthly subscriptionBased on token usage Pay-as-you-go
Scaling methodManually scale the provisioned throughputManually scale the number of model unitsSubmit a request for manual review in the console
Limitations- Daily subscriptions are prepaid and non-refundable. - If your usage exceeds the provisioned throughput, the system automatically routes requests to the Model Studio model call service.If you cancel a prepaid subscription within the first month, days used are billed at 1.2 times the standard daily rate.- Supports only specific models fine-tuned with LoRA. - The system automatically releases resources after one month of inactivity.

To view historical token usage and call counts, go to model monitoring.

Billing

Time-based billing (provisioned throughput)

Cost = Usage Duration × (Input TPM Unit Price × Input TPM + Output TPM Unit Price × Output TPM)

  • A subscription is activated upon payment and valid for N days, expiring at 23:59 on the Nth day. For orders placed after 22:00, the expiration date is extended by one day.

  • After a subscription expires, the service stops after a 2-hour grace period. Resources are retained for 14 hours before release.

  • You cannot terminate a subscription early.

  • For pay-as-you-go accounts with an overdue balance, deployed resources are retained and billed for 24 hours before automatic release.

If a model's input exceeds the max input tokens or the purchased TPM, the corresponding calls automatically switch to pay-as-you-go mode. In this case, inference performance may decrease, rate limiting reverts to the public traffic limits of the current snapshot model in your workspace, and fees are charged at the standard pay-as-you-go rate for model calls.

  • In this case, the API response header includes x-dashscope-ptu-overflow:true.

  • To view TPM statistics, go to model monitoring (China (Beijing)).

Singapore

Model nameModel codeThinking modeMax input tokensPay-as-you-go (Hourly)Subscription (Daily)
Input (per 10k TPM)Output (per 1k TPM)Input (per 10k TPM)Output (per 1k TPM)
DeepSeek-v3.2deepseek-v3.2Supported64,000$2.052$0.616$24.624$7.387
Qwen3.6-Plus-2026-04-02qwen3.6-plus-2026-04-02Supported128,000$1.200$0.720$14.400$8.640
Qwen3.5-plus-2026-04-20qwen3.5-plus-2026-04-20Supported128,000$0.960$0.576$11.520$6.912
Qwen3-vl-plus-2025-09-23qwen3-vl-plus-2025-09-23Supported128,000$0.480$0.384$5.760$4.608

China (Beijing)

Model nameModel codeThinking modeMax input tokensPay-as-you-go (Hourly)Subscription (Daily)
Input (per 10k TPM)Output (per 1k TPM)Input (per 10k TPM)Output (per 1k TPM)
GLM-5.1glm-5.1Supported64,000$2.97$1.19$35.65$14.26
Qwen3.6-Plus-2026-04-02qwen3.6-plus-2026-04-02Supported128,000$0.67$0.397$7.93$4.753
Qwen3.5-plus-2026-04-20qwen3.5-plus-2026-04-20Supported128,000$0.26$0.16$3.17$1.9
Qwen3-max-2025-09-23qwen3-max-2025-09-23Not supported128,000$1.11$0.45$13.32$5.40
Qwen-plus-2025-12-01qwen-plus-2025-12-01Not supported$0.28$0.07$3.36$0.84
Supported$0.28$3.36
Qwen-flash-2025-07-28qwen-flash-2025-07-28Supported$0.06$0.06$0.72$0.72
Qwen3-vl-plus-2025-09-23qwen3-vl-plus-2025-09-23Supported$0.35$0.35$4.20$4.20
DeepSeek-v3.2deepseek-v3.2Supported64,000$1.04$0.16$12.48$1.92

Pay-as-you-go (model unit)

Fee = Usage (hours) × Model units × Price per model unit

  • If you cancel a prepaid purchase within the first month, you are billed at 1.2 times the daily rate (any partial day is billed as a full day).

Note

Computing resources for pay-as-you-go model units are allocated on a first-come, first-served basis. If your purchase is unsuccessful, you receive a full refund.

Singapore

Model nameModel codeTypeRate limitingModel unitMax contextUnit price (Billed per minute; partial minutes are rounded up.)Monthly price (Billed per day; partial days are rounded up.) (For subscriptions canceled within the first month, used days are billed at 1.2 times the standard daily rate.)
Qwen3-32Bqwen3-32bInstructYesType I model unit (MU1)Fixed: 131,072$44/hour$20,916/month

Model type:

  • Instruct: Performs inference in non-thinking mode.

China (Beijing)

Model nameModel codeTypeRate limitingModel unitMax contextUnit price (Billed per minute; partial minutes are rounded up.)Monthly price (Billed per day; partial days are rounded up.) (For subscriptions canceled within the first month, used days are billed at 1.2 times the standard daily rate.)
Qwen3-14Bqwen3-14bInstructYesType I model unit (MU1)Fixed: 131,072$40/hour$18,800/month
Qwen3-32Bqwen3-32bInstructYesType I model unit (MU1)Fixed: 131,072$40/hour$18,800/month
Qwen3-VL-8B-Instructqwen3-vl-8b-instructInstructYesType I model unit (MU1)Fixed: 131,072$20/hour$9,400/month
Qwen3-VL-8B-Thinkingqwen3-vl-8b-thinkingThinkingYes

Model type:

  • Instruct: Performs inference in non-thinking mode.

  • Thinking: Performs inference in thinking mode.

Token-based billing

Cost = (Number of input tokens × Input unit price) + (Number of output tokens × Output unit price) (Minimum billing unit: 1 token)

  • Token-based billing applies only to custom models fine-tuned from the following foundation models using Supervised Fine-Tuning (SFT).

Singapore

Foundation modelModel codeModel typeMax contextInput unit priceOutput unit price
Qwen3-14Bqwen3-14b-instructInstructFixed at 131,072¥0.001/1,000 tokensNon-thinking mode: ¥0.004/1,000 tokens Thinking mode: ¥0.01/1,000 tokens

To deploy additional models, consult this solution to select the best deployment plan for your business.

Deployment method

To deploy a model in the console, follow these steps:

If you receive an "insufficient permissions" error, see Deployment permission errors.

- Go to the model deployment console (China (Beijing)).
- Select a model and billing method, leave other settings as default, specify a model name, and start deployment.
- A Running status indicates a successful deployment. * **Important ** Billing starts after the model deploys successfully.

Deployment configuration

Model unit

ParameterDescription
model inference modeWhen deploying certain models as a Model Unit, you can configure the inference mode and max context. - Instruct - Runs in non-thinking mode. - Thinking - Runs in thinking mode.
max contextThis setting is available for certain models deployed as a Model Unit. Max context length depends on the model type.
service throttlingThis setting is available for certain models deployed as a Model Unit. You can limit the requests per minute (RPM) and tokens per minute (TPM) for model calls.

Call a deployed model

After deploying a model, call it using the OpenAI-compatible API, DashScope, or the Assistant SDK.

When calling a deployed model, set the model parameter to the model code. Find the Model Code in the deployment console (China (Beijing)).

DashScope

HELPCODEESCAPE-python
import os
import dashscope

messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Who are you?"},
]
dashscope.base_http_api_url = 'https://dashscope-intl.aliyuncs.com/api/v1'
response = dashscope.Generation.call(
    # If the DASHSCOPE_API_KEY environment variable is not set, provide your Model Studio API key directly, for example: api_key="sk-xxx"
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    model="qwen3-max-xxx-xxx",  # Replace with the model code of your deployed model.
    messages=messages,
    result_format="message",
    enable_thinking=False,
)
print(response)

OpenAI-compatible API

HELPCODEESCAPE-python
import os
from openai import OpenAI


client = OpenAI(
    # If the DASHSCOPE_API_KEY environment variable is not set, provide your Model Studio API key directly, for example: api_key="sk-xxx"
    api_key=os.getenv('DASHSCOPE_API_KEY'),
    base_url="https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
    model="qwen3-max-xxx-xxx",  # Replace with the model code of your deployed model.
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Who are you?"},
    ],
    extra_body={"enable_thinking": False},
)
print(completion)

Service scaling

  • Provisioned throughput (time-based billing): Click Scaling to adjust the number of instances.

  • Model unit (time-based billing): Click Scaling to adjust the number of instances.

Service deactivation

Go to the model deployment console (China (Beijing)), find the deployment service to deactivate, and click Deactivate. After confirmation, billing for the service stops.

FAQ

Deploying your own models

Currently, you cannot upload and deploy your own models. Check Model Studio for the latest updates.

Alternatively, use Alibaba Cloud Platform for AI (PAI) to deploy your own models. For instructions, see Deploy large language models in PAI.

Deployment permission errors

  1. If you see the error message "You Do Not Have Permissions For This Module" , ensure your account has the ModelDeploy-FullAccess permission for the workspace.

    If you cannot proceed, contact your organization or IT administrator to grant the necessary permissions or investigate the issue.

  2. If you receive the error message "xx workspace does not have deployment privilege for model xx " during deployment, go to the Model Studio Workspaces page and authorize the workspace to deploy the model.

    API call error: Workspace xxx does not have deployment privilege for model xxxx.

    If you still receive permission errors, contact your organization or IT administrator to grant the necessary permissions or perform the operation on your behalf.

Mirror of Alibaba Cloud Model Studio docs for reference and RAG. Not affiliated with Alibaba Cloud.