Appearance
API Reference: Models
Browse documentation for api reference: models. Pages organized by topic:
API Usage
- Create an API key — Create an API key to call models and applications in Alibaba Cloud Model Studio
- Export an API key as an environment variable — Set your API key as an environment variable to avoid hard-coding it in source code and reduce the risk of accidental exposure
- Install the SDK — Model Studio provides official DashScope SDKs (Python, Java) and supports OpenAI SDKs (Python, Node
- Error messages — Lists the error messages that you may encounter when calling Alibaba Cloud Model Studio APIs, along with possible causes and solutions
CosyVoice Real
- time speech synthesis API reference-CosyVoice speech synthesis WebSocket API — Use a WebSocket connection to access the CosyVoice real-time speech synthesis service
- time speech synthesis API reference-CosyVoice client events — User guide: For model introduction and selection recommendations, see Speech synthesis
- time speech synthesis API reference-CosyVoice server-side events — User guide: For model introduction and selection recommendations, see Speech synthesis
- time speech synthesis API reference-CosyVoice speech synthesis Java SDK — The parameters and key interfaces of the CosyVoice speech synthesis Java SDK
- time speech synthesis API reference-CosyVoice speech synthesis Python SDK — The parameters and interfaces of the CosyVoice speech synthesis Python SDK
- time speech synthesis API reference-CosyVoice voice list — System voices for CosyVoice are listed below
Custom Hotwords API Reference
- Custom vocabulary HTTP API reference — Manage custom vocabularies through HTTP APIs, including creating, listing, getting, updating, and deleting vocabularies
- Custom hotwords Java SDK reference — Use the Java SDK to create, query, update, and delete custom vocabularies for speech recognition
- Custom hotwords Python SDK reference — Use the Python SDK to create, list, query, update, and delete custom vocabularies for speech recognition
Fine
- tuning (training)-Video generation model fine-tuning API reference — Train a domain-specific LoRA model for Wan image-to-video generation using REST APIs
Fun
- ASR real-time speech recognition API reference-Python SDK — The parameters and interfaces for the Fun-ASR real-time speech recognition Python SDK
- ASR real-time speech recognition API reference-Java SDK — The interfaces and parameters of the Fun-ASR Java SDK for real-time speech recognition
- ASR real-time speech recognition API reference-WebSocket API — Use the WebSocket protocol to integrate directly with the Fun-ASR real-time speech recognition service
Generate Dance Videos From Images
- AnimateAnyone-AnimateAnyone image detection API reference — Ensure images meet AnimateAnyone Gen2 requirements before triggering video generation
- AnimateAnyone-AnimateAnyone action template generation API reference — Extract character movements from motion videos using the animate-anyone-template-gen2 model
- AnimateAnyone-AnimateAnyone video generation API reference — Automate character action video production with the AnimateAnyone API
Image Generation
- Qwen-Image API reference — Qwen-Image is a general-purpose image generation model that supports multiple artistic styles and excels at complex text rendering
- Qwen-Image Edit API reference — Qwen-Image Edit supports multi-image input and output
- Qwen-MT-Image API reference — Qwen-MT-Image accurately translates text in images while preserving the original layout
- Z-Image API reference — Generate multilingual images via Z-Image API using a synchronous HTTP POST to dashscope-intl endpoints
- Wan text-to-image V2 API reference — The Wan text-to-image model generates images from text prompts, supporting artistic styles and realistic photographic effects
- Wan2.7 - image generation and editing — Wan2
- Wan2.6 - image generation and editing — The Wan2
- Wan2.5 - general image editing — Edits images using text prompts
- Wan2.1 - general image editing API reference — Implement inpainting, super resolution, colorization, and watermark removal using Wan2
- Image erase completion API reference — Eliminate unwanted people, watermarks, and objects from images with the image-erase-completion model
- FAQ — Understand how Model Studio's image generation API handles free quota billing, QPS rate limits, and BadRequest
Image To Broadcast Video
- LivePortrait-LivePortrait image detection API reference — Validate portrait images against LivePortrait specs before synthesis — use the face-detect HTTP POST endpoint to catch oversized files, unsu
- LivePortrait-LivePortrait video generation API reference — Automate portrait animation at scale using the LivePortrait API: submit async tasks via POST with X-DashScope-Async header, poll results by
Image To Emoji Video
- Emoji-Emoji image detection API reference — Ensure portrait images pass Emoji video pre-checks using emoji-detect-v1
- Emoji-Emoji video generation API reference — Generate facial Emoji videos from portrait images using the emoji-v1 model
Image To Singing Video
- EMO-EMO image detection API reference — Avoid failed EMO video generation by validating portrait images first
- EMO-EMO video generation API reference — Generate lip-synced talking portrait videos with the EMO API — submit a portrait image and voice audio via HTTP POST, then poll async result
Lip
- sync replacement for videos - VideoRetalk-VideoRetalk API reference — Produce realistic lip-synced speaking videos using the VideoRetalk API
More
- Generate a temporary API key — To call model services from untrusted environments such as browsers and mobile apps, use a secure backend service to generate temporary API
- Asynchronous task management API — Model Studio's asynchronous task APIs let you query individual task results by task_id, batch-check multiple task statuses, and cancel queue
- Call a model in a sub-workspace — Isolate model access and billing across teams using sub-workspace API keys
- Upload local files to get temporary URLs — Multimodal models such as Qwen-VL require file URLs for image, video, and audio inputs
- Configure connection reuse for DashScope SDK — Without connection reuse, each API call opens a new TCP connection and performs a TLS handshake, adding latency
More Models
- Text rerank — Retrieval systems often return imprecise results during initial retrieval
- Intent recognition — Detect user intents in milliseconds and route to the right tools automatically—tongyi-intent-detect-v3 with INTENT_MODE eliminates manual ro
Multimodal Embedding
- Multimodal embeddings API — Multimodal embedding models convert text, images, and videos into embeddings in a shared semantic space, enabling cross-modal retrieval, con
Music Generation
- Music generation API reference(Fun-Music) — API parameters for the Fun-Music music generation model
OpenAI
- compatible Responses-Create a response — Call Qwen models using the OpenAI-compatible Responses API
- compatible Responses-Retrieve a response — Retrieve a completed model response by its Response ID
- compatible Responses-Delete a response — Deletes a stored model response based on its response ID
- compatible Responses-List input items — Returns the input items used to generate the specified Response
OutfitAnyone
- AI Try-on Plus API reference — Compared to aitryon, aitryon-plus offers superior image clarity, fabric texture, and logo restoration
- OutfitAnyone-Parsing API reference — Enable accurate partial try-on by isolating garment regions with aitryon-parsing-v1
Paraformer Real
- time speech recognition API reference-Java SDK — The parameters and interfaces of the Paraformer real-time speech recognition Java SDK
- time speech recognition API reference-Python SDK — The parameters and interfaces of the Paraformer real-time speech recognition Python SDK
- time speech recognition API reference-WebSocket API — ImportantThis document applies only to the China (Beijing) region
- time speech recognition API reference-Performance optimization for real-time speech recognition in high-concurrency scenarios — Boost Paraformer real-time speech recognition performance using DashScope Java SDK connection pools and object pools
Paraformer Recording File Recognition API Reference
- Java SDK — The parameters and API of the Paraformer audio file recognition Java SDK
- Python SDK — The parameters and API of the Paraformer audio file recognition Python SDK
- RESTful API — The parameters and API details for Paraformer audio file recognition RESTful API
- Best practices — Optimize video file transcription for Paraformer with a single FFmpeg command to extract and compress audio, reducing file size and boosting
Real
- time multimodal-Client events — Client events for the Qwen-Omni-Realtime API
- time multimodal-Server events — Server events for the Qwen-Omni-Realtime API, including function calling events
- time multimodal-Python SDK — The key interfaces and request parameters for Qwen-Omni real-time using the DashScope Python SDK
- time multimodal-Java SDK — The key interfaces and request parameters for Qwen-Omni real-time DashScope Java SDK
- time multimodal-Real-time multimodal interaction flow — The interaction flow between the real-time multimodal server and client
- time multimodal-Voice cloning API reference — Voice cloning lets you clone voices without training
- time speech synthesis API reference (Qwen-TTS-Realtime)-Interaction flow for real-time speech synthesis — Connect to the Qwen-TTS real-time speech synthesis service over a WebSocket connection
- time speech synthesis API reference (Qwen-TTS-Realtime)-Client events — Understand the client events for the Qwen-TTS Realtime API to precisely control the synthesis lifecycle, from session updates to text buffer
- time speech synthesis API reference (Qwen-TTS-Realtime)-Server events — Learn the server events for the Qwen-TTS-Realtime API, from session creation to completion
- time speech synthesis API reference (Qwen-TTS-Realtime)-Python SDK — This topic describes the key interfaces and request parameters for calling real-time speech synthesis (Qwen) with the DashScope Python SDK
- time speech synthesis API reference (Qwen-TTS-Realtime)-Java SDK — Key interfaces and request parameters for Qwen real-time speech synthesis DashScope Java SDK
- time speech recognition (Qwen-ASR-Realtime)-Client events for Qwen-ASR-Realtime — Master Qwen-ASR Realtime client events to maintain low-latency WebSocket sessions and reliable multilingual transcription
- time speech recognition (Qwen-ASR-Realtime)-Server events for Qwen-ASR-Realtime — Master Qwen-ASR-Realtime's WebSocket event lifecycle: from session
- time speech recognition (Qwen-ASR-Realtime)-Qwen-ASR-Realtime Python SDK - API reference — Build a production-grade real-time speech recognition pipeline using DashScope Python SDK
- time speech recognition (Qwen-ASR-Realtime)-Qwen-ASR-Realtime Java SDK - API reference — Build real-time speech recognition into your Java app using DashScope SDK
- time speech recognition (Qwen-ASR-Realtime)-Interaction flow for Qwen-ASR-Realtime — Qwen-ASR-Realtime leverages WebSocket to deliver low-latency speech recognition, offering VAD mode for automatic silence detection and Manua
- time audio and video translation - Qwen-Livetranslate-Realtime-Client events — Master the WebSocket client event schema for qwen3-livetranslate-flash-realtime: session
- time audio and video translation - Qwen-Livetranslate-Realtime-Server events — Build reliable real-time translation pipelines by mastering qwen3-livetranslate-flash-realtime server events
- time audio and video translation - Qwen-Livetranslate-Realtime-Qwen-LiveTranslate Python SDK - API reference — Build real-time multilingual speech translation with the DashScope Python SDK
- time audio and video translation - Qwen-Livetranslate-Realtime-Qwen-LiveTranslate Java SDK - API reference — Integrate real-time speech translation into Java apps using DashScope SDK v2
Recording File Recognition API Reference (Fun
- ASR)-Python SDK — Parameters and interface details for the Fun-ASR audio file recognition Python SDK
- ASR)-RESTful API — The parameters and API details of the Fun-ASR RESTful API for audio file recognition
- ASR)-Java SDK — The parameters and interface details of the Fun-ASR audio file recognition Java SDK
Speech Recognition
- Audio file recognition (Qwen-ASR) API reference — Input and output parameters for the Qwen-ASR model
Speech Synthesis
- Speech synthesis (Qwen-TTS) — Request parameters and response fields for the Qwen-TTS speech synthesis API
- Voice Design API reference — Use the Voice Design HTTP API to create, list, query, and delete custom voices
Speech Translation
- Audio and video translation - Qwen API reference — qwen3-livetranslate-flash delivers real-time audio and video translation via an OpenAI-compatible chat completions endpoint
Text Embedding
- Synchronous API details — Explore the synchronous API for text embedding, detailing OpenAI-compatible and native DashScope interfaces to convert text into vectors for
Text Generation API Reference
- OpenAI-compatible Chat — Call models using the OpenAI-compatible Chat API, including input and output parameter descriptions and code examples
- Anthropic-compatible Messages — Call models via the Anthropic-compatible Messages API
- DashScope API reference — Call models using the DashScope API
Text Translation
- Qwen-MT API reference — The input and output parameters for calling Qwen-MT through the OpenAI compatible interface or the DashScope API
- Qwen-Deep-Research API reference — Request and response parameters for calling the Qwen-Deep-Research model through the DashScope API
- Qwen-OCR API reference — Extract text from invoices and documents using Qwen-VL-OCR via OpenAI-compatible and DashScope APIs
Toolkit/Framework
- OpenAI compatible - Chat — The Qwen model in Alibaba Cloud Model Studio supports an OpenAI-compatible interface
- OpenAI compatible - Responses — Alibaba Cloud Model Studio supports the OpenAI-compatible Responses API
- OpenAI compatible - Completions — Qwen2
- OpenAI Vision interface compatibility — Alibaba Cloud Model Studio’s Qwen vision models are compatible with the OpenAI API specification
- OpenAI compatible file interface — Streamline file lifecycle management on Model Studio using the OpenAI-compatible API
- OpenAI compatible - Batch Chat — For non-real-time scenarios like data annotation and content generation, the Batch Chat API offers a low-cost, high-concurrency alternative
- OpenAI compatible - Embedding — Alibaba Cloud Model Studio's embedding models are compatible with the OpenAI API
- OpenAI compatible - Conversations — Manually managing message lists for conversations that span multiple devices or have long interruptions can lead to context loss
Video Generation
- HappyHorse text-to-video API reference — HappyHorse text-to-video generates physically realistic, motion-smooth video from text prompts
- HappyHorse - image-to-video-first frame API reference — The HappyHorse image-to-video model generates physically realistic and smooth-motion videos from a first frame and a text prompt
- HappyHorse - reference-to-video API reference — The HappyHorse reference-to-video model lets you provide multiple reference images and a text prompt to generate a video that combines subje
- HappyHorse video editing API reference — The HappyHorse video editing model takes a video and a reference image as input, and performs editing tasks such as style transfer and local
- Wan 2.7 - image-to-video API — The Wan 2
- Wan - text-to-video API reference — The Wan text-to-video model generates smooth videos from text prompts
- Wan - reference-to-video API reference — Wan-R2V supports multimodal input
- Wan2.7 - video editing API reference — The Wan 2
- Wan image-to-action API reference — Animate a character image by transferring actions from a reference video
- Wan - video character swap API reference — Replaces the main character in a video with a character from an image while preserving the original scene, lighting, and tone for seamless i
- Video style transform API Reference — Automate video stylization pipelines with step-by-step API integration: send asynchronous POST requests to the video-synthesis endpoint, pol
Voice Cloning API Reference
- Voice cloning HTTP API reference — Use the HTTP API to create, list, query, update, and delete cloned voices
- Voice cloning Java SDK reference — Use the DashScope Java SDK to clone and manage CosyVoice voices
- Voice cloning Python SDK reference — CosyVoice voice cloning is accessible through the DashScope Python SDK
Wan
- image-to-video - first frame API reference(2.1-2.6)-Wan first frame to video - Video effect templates — Generate dynamic videos from a first-frame image using Wan video effect templates like 'flying' or 'squish'
- legacy video models-Wan - text-to-video API reference (2.1-2.6) — The Wan text-to-video model generates smooth videos from text prompts
- legacy video models-Wan - reference-to-video (2.6) — The Wan reference-to-video model accepts multimodal input and generates single-character or multi-character interaction videos using people
- legacy video models-Wan -image-to-video-first and last frames API reference(2.2) — The Wan 2
- legacy video models-Wan - video editing (2.1) — The Wan 2
- digital human-wan2.2-s2v image detection API reference — Prevent failed Wan2
- digital human-wan2.2-s2v API reference — Generate portrait lip-sync videos from a single image and audio file using Wan2