Appearance
User Guide: Models
Browse documentation for user guide: models. Pages organized by topic:
Get Started
- What is Alibaba Cloud Model Studio — Alibaba Cloud Model Studio is a one-stop model service platform
- Make your first API call to Qwen — This guide walks you through making your first API call to a Qwen model on Alibaba Cloud Model Studio
- Recommended models — Alibaba Cloud Model Studio offers Qwen and third-party models for text, image, audio, and video
- Rate limits — Model Studio enforces rate limits to ensure fair use across all RAM users, workspaces, and API keys under one Alibaba Cloud account
Billing
- Free quota for new users — Activate Model Studio in the Singapore region to receive free quota for each model
- Model inference pricing — Tiered pricing rulesSome Model Studio models use tiered pricing
- Training and deployment pricing — This topic describes the billing rules and pricing for model training and model deployment on Alibaba Cloud Model Studio
- Savings plans — Model Studio offers savings plans and resource plans to help you reduce model costs
- Bill query and cost management — Query billing details, analyze bills, and stop billing
Text Generation
- Overview — Text generation models produce clear and coherent text based on input prompts
- Multi-turn conversations — Qwen API is stateless — pass conversation history via a messages array each request
- Streaming output — Stream Qwen model responses token by token using Server-Sent Events over persistent HTTP connections
- Deep thinking — Deep thinking models reason before responding, which improves accuracy on complex tasks such as logical reasoning and math
- Structured output — Structured output (JSON mode) makes the model return a valid JSON string that your code can parse directly, without extra text like json tha
- Partial mode — For scenarios like code completion and text continuation, you can generate new content starting from an existing text fragment (prefix)
- Context cache — Inference requests to a large model often include overlapping input, such as in a multi-turn conversation or a series of questions about the
- Batch inference — Process large volumes of requests asynchronously at 50% the cost of real-time inference
Tool Calling
- Function calling — Large language models may struggle with tasks that require real-time data or complex math
- Web search — Large language models (LLMs) rely on training data that has a knowledge cutoff date
- Web extractor — LLMs cannot directly access web page data
- Code Interpreter — Enable the built-in Python Code Interpreter when calling a model
- Text-to-image search — The text-to-image search tool enables a model to search the Internet for relevant images based on a text description
- Search by image — The search by image tool enables the model to search the Internet for visually similar images based on an input image
- Knowledge retrieval — Large Language Models (LLMs) cannot answer questions about private data
- MCP — The Model Context Protocol (MCP) enables large language models to use external tools and data
Best Practices
- Integrate tools — Token Plan (Team Edition) supports extending AI coding tool capabilities through built-in model tools, including web search, code interprete
- Integrate multimodal generation models — Image generation models must be integrated through each tool's extension mechanism (Skill, Slash Command, or Agent)
Best Practices
- Text-to-text prompt guide — A prompt is the text you input to a large language model (LLM), used to explicitly tell the model what problem you want to solve or what tas
- Text-to-image prompt guide — Eliminate vague AI outputs with structured prompt formulas for Wan V2 text-to-image
- Text-to-video/image-to-video prompt guide — Use structured prompt formulas to control content, motion, camera work, audio, and visual style in AI-generated videos
- Best practices for handling rate limiting — Model Studio APIs limit request volume, token usage, and growth rate
Changelog
- Model lifecycle and updates — The following tables list model releases
- Announcements — Stay current with Alibaba Cloud Model Studio updates: Qwen3-235B-A22B price cuts, Qwen3 model family releases, batch inference API support,
Clients And Developer Tools
- OpenClaw — OpenClaw is an open source personal AI assistant platform that lets you interact with AI through various messaging channels
- Hermes Agent — Hermes Agent is a terminal AI coding tool
- Claude Code — Claude Code is a command-line AI coding assistant developed by Anthropic
- OpenCode — OpenCode is a terminal AI coding tool that connects to Alibaba Cloud Model Studio via Pay-as-you-go, Coding Plan, or Token Plan (Team Editio
- Cursor — Cursor is an AI coding IDE
- Codex — Codex is a terminal AI coding assistant developed by OpenAI
- Qwen Code — Qwen Code is a terminal-based AI coding tool that can be connected to Alibaba Cloud Model Studio via pay-as-you-go, Coding Plan, or Token Pl
- Cherry Studio — Cherry Studio is an open-source AI desktop client
- Chatbox — Chatbox is a cross-platform AI client that connects to Model Studio using Token Plan (Team Edition), Coding Plan, or Pay-as-you-go
- Cline — Cline is a VS Code extension for AI-assisted coding
- Qoder — Qoder is an agentic coding platform for software development that supports a desktop IDE, CLI, and JetBrains plugin, and can connect to Alib
- Lingma — Lingma is an Alibaba Cloud smart coding assistant that provides a standalone IDE and can connect to Alibaba Cloud Model Studio via Coding Pl
- Kilo CLI — Kilo CLI is the command-line client for Kilo Code
- Call image and video generation APIs with Postman and cURL — Use Postman or cURL to call image or video generation APIs in Model Studio
- Dify — Dify is an open-source platform for building AI applications
- More tools — In addition to the tools listed above, Alibaba Cloud Model Studio supports any third-party programming tool that is compatible with the Open
Coding Plan
- Coding plan overview — Coding Plan gives you access to models including Qwen, GLM, Kimi, and MiniMax in popular AI coding tools
- web search — Add a web search tool to Coding Plan so the model can retrieve real-time information
- add visual understanding capabilities — Some models in Model Studio Coding Plan, such as qwen3
- FAQ — Connection and configurationCommon errors and solutionsError messagePossible causeSolution400 InvalidParameter: Range of input length should
Deployment
- Model deployment — Deploying a provides a dedicated inference service that supports performance requirements such as high concurrency and low latency
Embedding And Rerank
- Embedding — Convert text, images, and videos into dense or sparse vectors using Model Studio embedding models
- Rerank — Retrieval systems prioritize speed, which can sometimes compromise the precision of search results
Fine
- tuning-Fine-tune Qwen — Fine-tune Qwen in Alibaba Cloud Model Studio using an HTTP API
- tuning-Fine-tune a video generation model — When off-the-shelf Wan models lack your brand's visual style, fine-tune them on Model Studio using custom video datasets in JSONL format
Image Editing
- Qwen-Image-Edit — Qwen-Image-Edit supports multi-image input and output
- Image editing - Wan2.5 to 2.7 — The Wan image editing model series supports multi-image input and output
- General image editing - Wan2.1 — The Wan image editing model performs various editing tasks from text prompts: outpainting, watermark removal, style transfer, instruction-ba
Image Generation And Editing
- Text-to-image — Use the text-to-image API to generate images from text descriptions
Inference
- Music generation — Fun-Music generates full-length songs with male or female vocals in Chinese or English
- Speech-to-speech — Select a model for use cases like conversational speech and speech translation
Model Usage Statistics And Performance Monitoring
- Model usage — Understand how Model Studio measures AI consumption across tokens, image zhang, and video seconds — and how free-tier quotas work before HTT
Omni
- modal-Qwen-Omni-Realtime — Qwen-Omni-Realtime is a real-time audio and video chat model that processes streaming audio and image inputs, including video frames, to gen
- modal-Qwen-Omni — The Qwen-Omni model accepts multimodal input and generates text or speech responses
- modal-Real-time audio and video translation - Qwen — qwen3-livetranslate-flash-realtime is a vision-enhanced real-time translation model that translates between 18 languages, including Chinese,
- modal-Audio and Video File Translation – Qwen — Build a real-time multilingual translation pipeline with Qwen3-LiveTranslate-Flash
- modal-Voice list — Voices supported by Qwen-Omni (non-real-time) and Qwen-Omni-Realtime models, with their corresponding voice parameter values
Security And Compliance
- Permission management — Alibaba Cloud Model Studio offers granular access control at the console and model level, helping organizations manage users across multiple
- Security certifications and privacy — Compliance certificationsAlibaba Cloud Model Studio has achieved SOC 2 compliance with an unqualified opinion, validating its controls for S
- Alibaba Cloud Model Studio - Training data summary — Understand how Qwen and Wan models are built on multimodal corpora spanning text, images, video, and audio—with quality filters, deduplicati
Specialized Models
- Long context (Qwen-Long) — Qwen-Long handles documents up to 10 million tokens through a file upload and reference mechanism, overcoming standard model context limits
- Code Capabilities (Qwen-Coder) — Qwen-Coder is a language model designed for code-related tasks
- Machine translation (Qwen-MT) — Qwen-MT is a machine translation model fine-tuned from Qwen3
- Role-playing (Qwen-Character) — Qwen's role-playing model enables human-like conversations for virtual social apps, game non-player characters (NPCs), IP replication, and s
- Data mining (Qwen-Doc-Turbo) — The data mining model extracts information, moderates content, classifies data, and generates summaries
- Deep research (Qwen-Deep-Research) — Automates complex research through planning, multiple rounds of web searches, and structured report generation
- Mathematical capabilities (Qwen-Math) — Qwen-Math provides powerful mathematical reasoning and calculation capabilities, with detailed problem-solving steps for easy understanding
- Tongyi Xiaomi Conversation Analysis — Tongyi Xiaomi Conversation Analysis focuses on analytical tasks such as information extraction, scenario classification, and satisfaction as
- Audio understanding (Qwen3-Omni-Captioner) — Automate audio captioning with Qwen3-Omni-Captioner — no prompts needed
Speech
- to-text-Real-time speech recognition - Fun-ASR/Paraformer — The real-time speech recognition service converts an audio stream into punctuated text for a "text as you speak" experience
- to-text-Real-time speech recognition - Qwen — In scenarios such as live streaming, online meetings, voice chat, or smart assistants, you need to convert continuous audio streams into tex
- to-text-Audio file recognition - Fun-ASR/Paraformer — The Fun-ASR/Paraformer audio file recognition models convert recorded audio into text
- to-text-Audio file recognition - Qwen — Qwen audio file recognition models convert recorded audio to text
- to-text-Custom hotwords — Define a custom vocabulary list to improve the recognition accuracy of domain-specific terms, product names, and specialized vocabulary
Speech Synthesis
- Real-time speech synthesis — Real-time speech synthesis converts text into natural speech over a WebSocket connection
- Non-real-time speech synthesis — Non-real-time speech synthesis converts text to speech (TTS) through an HTTP API
- Voice cloning — Voice Cloning creates a highly realistic custom voice from a 10- to 20-second audio sample, with no model training required
- Voice Design — Voice Design lets you create custom voices from natural language descriptions alone, with no audio samples required
- SSML — Use SSML (Speech Synthesis Markup Language) to fine-tune speech characteristics such as speed, pauses, and pronunciation
Support
- FAQ — Frequently asked questions about Alibaba Cloud Model Studio
- Related agreements — Understand the legal and service commitments governing Alibaba Cloud Model Studio — covering Terms of Service, Model Inference SLA uptime gu
Third
- party model integration tutorial-DeepSeek — This topic describes how to call DeepSeek models on Alibaba Cloud Model Studio through the OpenAI-compatible API or the DashScope SDK
- party model integration tutorial-Kimi — Call Kimi models deployed on Alibaba Cloud Model Studio
- party model integration tutorial-GLM — Call GLM models on Alibaba Cloud Model Studio
- party model integration tutorial-MiniMax — Call MiniMax models on Alibaba Cloud Model Studio
Token Plan (Team Edition)
- Token Plan overview — Token Plan (Team Edition) is an AI model subscription from Alibaba Cloud Model Studio
- Quick start — Subscribe to Token Plan (Team Edition) and get started in three steps: choose a plan, obtain your API key, and configure your AI tool
- FAQ — Frequently asked questions about Token Plan (Team Edition), covering purchasing, usage, metering, and performance
Transmission Security
- Access Model Studio APIs over a private network by using an endpoint — To call Model Studio APIs from a VPC without exposing traffic to the public internet, create a private endpoint
Video Editing
- Video Editing 2.7 — The Wanxiang video editing model supports editing operations on input videos, such as adding, deleting, or modifying content, replacing back
- General video editing — The Wan unified video editing model supports multimodal input (text, image, and video) and provides five core capabilities: multi-image refe
Video Generation And Editing
- Text-to-video — Wan I2V supports multi-modal input (text, images, and audio) and generates videos up to 15 seconds long at 1080P resolution
- Image-to-video 2.7 — Wan2
- Image-to-video: first and last frames — The Wan image-to-video model generates smooth videos from a first-frame image, a last-frame image, and an optional text prompt
- Reference-to-video — Wan-R2V accepts multimodal input (text, image, video, and audio) to generate performance videos
Visual Understanding
- Image and video understanding — Visual understanding models can answer questions based on the images or videos that you provide
- Text extraction (Qwen-OCR) — Automate document digitization with Qwen-OCR — extract structured text from scanned files, tables, and receipts via OpenAI-compatible and Da
- Visual reasoning — Visual reasoning models output their thinking process before answering