Appearance
Model lifecycle and updates
The following tables list model releases . For model deprecation rules and lists, see Model deprecation.
Global
If you select the Global deployment scope, model inference compute resources are dynamically scheduled worldwide. Static data is stored in your selected region. Supported regions: US (Virginia) and Germany (Frankfurt).
| Type | Time | Model | Description |
|---|---|---|---|
| Reasoning model | 2026-05-20 | glm-5.1 | The Zhipu GLM-5.1 model is designed for long-horizon tasks and supports a 200K context window with a maximum output of 128K tokens. With strong logical reasoning, long-text comprehension, and code generation capabilities, it delivers excellent results across multiple benchmarks and is well suited for intelligent interaction, enterprise applications, and development assistance. GLM-Alibaba Cloud |
| Text-to-video | 2026-05-06 | happyhorse-1.0-t2v | HappyHorse text-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - text-to-video |
| Image-to-video | 2026-05-06 | happyhorse-1.0-i2v | HappyHorse image-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - image-to-video-first frame |
| Video editing | 2026-05-06 | happyhorse-1.0-video-edit | HappyHorse video editing model. Supports video editing and processing. HappyHorse - video editing |
| Reference-to-video | 2026-05-06 | happyhorse-1.0-r2v | HappyHorse reference-to-video model. Supports multiple reference images as input to generate audio videos with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - reference-to-video |
| Text-to-text \& image understanding | 2026-04-29 | kimi-k2.5 | A visual understanding model from Moonshot AI that excels in general-purpose intelligent tasks such as code generation and visual understanding. It also supports image, video, and text input, along with dialogue and Agent tasks. Kimi on Alibaba Cloud |
| Reasoning model | 2026-04-29 | deepseek-v4-pro | The unit price of cached_token for deepseek-v4-pro is updated to USD 0.13752131580395 per million tokens. The standard input_token price remains unchanged. See Context cache. |
| Reasoning model | 2026-04-02 | qwen3.6-plus, qwen3.6-plus-2026-04-02 | Qwen3.6-Plus features significant upgrades to its code development capabilities, such as Agentic Coding and frontend programming, and provides a significantly enhanced Vibe Coding experience. Its inference capabilities in generalization scenarios are further enhanced. For multimodal capabilities, its performance in universal object recognition, OCR, and object localization is significantly improved. This version also fixes known issues that were identified after Qwen3.5-Plus was published. The usage method is the same as that of qwen3.5-plus. Overview |
| Reasoning model | 2026-03-04 | qwen3.5-flash, qwen3.5-flash-2026-02-23, qwen3.5-122b-a10b, qwen3.5-27b, qwen3.5-35b-a3b | The latest Qwen3.5-Flash model and open-source model from Alibaba, supporting text, image, and video input with fast response speed. Its overall performance is close to qwen3.5-plus. Supports built-in Text generation overview. Overview |
| Reasoning model | 2026-03-04 | qwen3.5-plus, qwen3.5-plus-2026-02-15, qwen3.5-397b-a17b | The latest Qwen3.5-Plus model and open-source model from Alibaba, supporting text, image, and video input. It delivers outstanding performance across language understanding, logical reasoning, code generation, agent tasks, image understanding, video understanding, GUI, and other tasks. Supports built-in Text generation overview. Overview |
| Image generation and editing | 2026-01-04 | wan2.6-image | Supports image editing and mixed image-text output. Wan - image generation and editing 2.6 |
| Image-to-video - first frame | 2026-01-04 | wan2.6-i2v | Adds the multi-shot narrative feature, which supports audio using automatic dubbing or custom audio files. Wan - image-to-video - first frame-based |
| Reference video | 2026-01-04 | wan2.6-r2v | Generates a multi-shot video based on a character's appearance and voice from a reference video. It also supports automatic dubbing. Wan2.7 - reference-to-video |
| Text-to-video | 2026-01-04 | wan2.6-t2v | Adds the multi-shot narrative feature, which supports audio using automatic dubbing or custom audio files. Wan - text-to-video |
| Visual understanding | 2026-01-04 | qwen3-vl-flash, qwen3-vl-flash-2025-10-15 | This small-size visual understanding model from the Qwen3 series effectively integrates thinking and non-thinking modes. Compared to the open-source Qwen3-VL-30B-A3B, it delivers better performance and a faster response time. Image and video understanding |
| Visual understanding | 2026-01-04 | qwen3-vl-8b-thinking, qwen3-vl-8b-instruct | An 8B dense open-source model from the Qwen3-VL series, available in both thinking and non-thinking versions. It uses less GPU memory and performs multimodal understanding and inference. It supports ultra-long contexts such as long videos and documents, 2D/3D visual positioning, and comprehensive spatial intelligence and object recognition. Image and video understanding |
| Visual understanding | 2026-01-04 | qwen3-vl-32b-thinking, qwen3-vl-32b-instruct | A 32B Dense model from the Qwen3-VL series. Its overall performance is second only to the Qwen3-VL-235B model. It excels in document recognition and understanding, spatial intelligence, object recognition, visual 2D detection, and spatial reasoning. This makes it suitable for complex perception tasks in general scenarios. Image and video understanding |
| Visual understanding | 2026-01-04 | qwen3-vl-plus, qwen3-vl-plus-2025-09-23, qwen3-vl-235b-a22b-thinking, qwen3-vl-235b-a22b-instruct | A visual understanding model from the Qwen3 series. It effectively combines thinking and non-thinking modes, and its visual agent capabilities are among the best in the world. This version features comprehensive upgrades in visual encoding, spatial intelligence, and multimodal thinking. It also significantly improves visual perception and recognition capabilities. Image and video understanding |
| Reasoning model | 2026-01-04 | qwen3-next-80b-a3b-thinking, qwen3-next-80b-a3b-instruct | A next-generation open-source model based on Qwen3. The thinking modelhas improved instruction following and more concise summaries compared with qwen3-235b-a22b-thinking-2507. See Deep thinking.The instruct modelhas enhanced Chinese comprehension, logical reasoning, and text generation compared with qwen3-235b-a22b-instruct-2507. See Overview. |
| Reasoning model | 2026-01-04 | qwen3-max, qwen3-max-2025-09-23 | Compared to the qwen3-max-preview version, this model has been specifically upgraded for agent programming and tool calling. This official release achieves state-of-the-art (SOTA) performance in its domain and supports more complex agent requirements. |
| Reasoning model | 2026-01-04 | qwen3-max-preview | The Qwen-Max model is a preview version based on Qwen3. Compared to the Qwen 2.5 series, its general-purpose capabilities are greatly improved, delivering significantly enhanced performance in Chinese and English text comprehension, complex instruction following, subjective open-ended tasks, multilingual processing, and tool calling. The model also produces fewer knowledge hallucinations. Qwen-Max |
| Code model | 2026-01-04 | qwen3-coder-flash, qwen3-coder-flash-2025-07-28 | The fastest and most cost-effective model in the Qwen Coder series. Code Capabilities (Qwen-Coder). |
| Code model | 2026-01-04 | qwen3-coder-plus, qwen3-coder-plus-2025-07-22, qwen3-coder-30b-a3b-instruct, qwen3-coder-480b-a35b-instruct | A code generation model based on Qwen3. It has strong Coding Agent capabilities and excels at tool calling and environment interaction. It combines excellent code capabilities with general-purpose abilities. Code Capabilities (Qwen-Coder) |
| Reasoning model | 2026-01-04 | qwen3-30b-a3b-thinking-2507, qwen3-30b-a3b-instruct-2507 | An upgrade of qwen3-30b-a3b. The thinking modelhas improved logical, general, knowledge, and creative capabilities. See Deep thinking.The instruct modelhas improved creative capabilities and model safety. See Overview. |
| Reasoning model | 2026-01-04 | qwen3-235b-a22b-thinking-2507, qwen3-235b-a22b-instruct-2507 | An upgrade of qwen3-235b-a22b. The thinking modelhas significantly improved logical, general, knowledge, and creative capabilities, suitable for challenging reasoning scenarios. See Deep thinking.The instruct modelhas improved creative capabilities and model safety. See Overview. |
| Reasoning model | 2026-01-04 | qwen3-30b-a3b, qwen3-32b, qwen3-14b, qwen3-8b | Qwen3 models support both thinking and non-thinking modes. You can switch between the two modes by using the enable_thinking parameter. In addition, the capabilities of Qwen3 models have been greatly enhanced: - 1. Reasoning capability: In evaluations of mathematics, code, and logical reasoning, it significantly outperforms QwQ and non-reasoning models of the same size, and achieves top-tier performance among industry models of a comparable scale. - 2. Human preference alignment: Capabilities for creative writing, role-playing, multi-turn conversation, and instruction following are greatly improved. The general capabilities significantly exceed those of models of a similar size. - 3. Agent capability: The agent achieves industry-leading performance in both inference and non-inference modes. It accurately invokes external tools. - 4. Multilingual capability: The models support over 100 languages and dialects and have significantly improved capabilities in multilingual translation, instruction understanding, and common-sense reasoning. - 5. Response format fixes: This version fixes response format issues from previous versions, such as abnormal Markdown, mid-sentence truncation, and incorrect boxed output. For more information about thinking mode, see Deep thinking. For more information about non-thinking mode, see Overview. |
| Text extraction | 2026-01-04 | qwen-vl-ocr-2025-11-20 | This snapshot of the Qwen text extraction model is based on the Qwen3-VL architecture and significantly improves document parsing and text localization. Text extraction (Qwen-OCR) |
| Text extraction | 2026-01-04 | qwen-vl-ocr | qwen-vl-ocr is a model specialized for OCR, with significantly improved text extraction from tables, exam papers, and similar images. For details, see . |
| Reasoning model | 2026-01-04 | qwen-plus-2025-12-01, qwen-plus-2025-09-11 | A model from the Qwen3 series. Compared to qwen-plus-2025-07-28, it has improved instruction-following capabilities and provides more concise summaries in thinking mode. See Deep thinking. In non-thinking mode, it has enhanced Chinese language comprehension and logical reasoning capabilities. See Overview. |
| Reasoning model | 2026-01-04 | qwen-plus-2025-07-28 | A model from the Qwen3 series. Compared to the previous version, it increases the context length to 1,000,000. For more information about thinking mode, see Deep thinking. For more information about non-thinking mode, see Overview. |
| Reasoning model | 2026-01-04 | qwen-plus | Offers a balance of capabilities. Its inference performance, cost, and speed are between those of Qwen-Max and Qwen-Flash, making it suitable for moderately complex tasks. |
| Multilingual translation | 2026-01-04 | qwen-mt-lite | A basic text translation model from Qwen. It supports translation between 31 languages. It offers a faster response and lower cost than qwen-mt-flash, making it suitable for latency-sensitive scenarios. Machine translation (Qwen-MT) |
| Multilingual translation | 2026-01-04 | qwen-mt-plus, qwen-mt-flash | The Qwen-MT model is a large language model for machine translation, optimized from the Qwen model. It excels at translation between Chinese and English, between Chinese and minor languages, and between English and minor languages. It supports 26 minor languages, such as Japanese, Korean, French, Spanish, German, Portuguese (Brazil), Thai, Indonesian, Vietnamese, and Arabic. In addition to multilingual translation, it offers features such as terminology intervention, domain prompting, and translation memory to improve translation quality in complex scenarios. See Machine translation (Qwen-MT). |
| Text-to-text | 2026-01-04 | qwen-flash, qwen-flash-2025-07-28 | The fastest and most cost-effective model in the Qwen series, suitable for simple jobs. |
International
If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.
| Type | Time | Model | Description |
|---|---|---|---|
| Text-to-video | 2026-04-27 | happyhorse-1.0-t2v | HappyHorse text-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - text-to-video |
| Image-to-video | 2026-04-27 | happyhorse-1.0-i2v | HappyHorse image-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - image-to-video-first frame |
| Reference-to-video | 2026-04-27 | happyhorse-1.0-r2v | HappyHorse reference-to-video model. Supports multiple reference images as input to generate audio videos with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - reference-to-video |
| Video editing | 2026-04-27 | happyhorse-1.0-video-edit | HappyHorse video editing model. Supports video editing and processing. HappyHorse - video editing |
| Text-to-video | 2026-04-26 | wan2.7-t2v-2026-04-25 | A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same as wan2.7-t2v. Wan - text-to-video |
| Image-to-video | 2026-04-26 | wan2.7-i2v-2026-04-25 | A snapshot version of the Wan 2.7 image-to-video model. The model capabilities are the same as wan2.7-i2v. Wan2.7 - image-to-video |
| Image generation | 2026-04-23 | qwen-image-2.0-pro-2026-04-22 | A model in the Qwen-Image-2.0 series that unifies image generation and editing. Compared with the March 3 snapshot, this model delivers a significant leap in visual quality --- especially texture detail, lighting, and materials. It supports multilingual in-image text generation and delivers more balanced artistic style performance. |
| Reasoning model | 2026-04-23 | qwen3.5-plus-2026-04-20 | A new snapshot of the Qwen3.5 native vision-language Plus model. Compared with the February 15 snapshot, agentic coding capabilities are significantly improved and inference speed is notably faster. Knowledge, reasoning, and long-context capabilities remain at a high level, making it well suited for agentic coding, production workflows, and high-throughput scenarios. |
| Reasoning model | 2026-04-23 | qwen3.6-27b | A 27B native vision-language dense model in the Qwen3.6 series. Compared with Qwen3.5-27B, agentic coding capabilities are significantly improved, and STEM and reasoning abilities are further enhanced. On the vision side, spatial intelligence and object localization & detection are notably stronger, while video understanding, document OCR, and visual Agent capabilities are steadily improved. |
| Reasoning model | 2026-04-20 | qwen3.6-max-preview | The largest proprietary model in the Qwen3.6 series, with further improved coding capabilities and more efficient agent execution. Supports text-only input, thinking mode (enabled by default), explicit caching, and function calling. Overview ** Image and video inputs are not supported. |
| Reasoning model | 2026-04-16 | qwen3.6-flash, qwen3.6-flash-2026-04-16, qwen3.6-35b-a3b | A native vision-language Flash model in the Qwen3.6 series, delivering significant overall performance improvements over Qwen3.5-Flash. Agentic coding capabilities are substantially enhanced (far surpassing the predecessor on multiple code agent benchmarks), along with stronger math and code reasoning. On the vision side, spatial intelligence is notably improved, with object localization and detection being particularly outstanding. Overview |
| Speech recognition | 2026-04-08 | fun-asr, fun-asr-2025-11-07 | Fun-ASR real-time speech recognition capability upgrades: - Dialect support: Covers the seven major Chinese dialects (Mandarin, Wu, Xiang, Gan, Hakka, Min, Yue) and 20+ regional accents - Ancient poetry optimization: Improves ancient poetry recognition accuracy for cultural education and audiobook scenarios - Text optimization: Enhances punctuation prediction and text normalization. Numbers, dates, and amounts automatically convert to standard formats - Multi-language expansion: Supports 30 languages, including Chinese, English, Japanese, Korean, French, German, and Spanish For more information, see Audio file recognition - Fun-ASR/Paraformer. |
| Video editing | 2026-04-03 | wan2.7-videoedit | The Wan 2.7 video editing model supports instruction-based editing and video transfer tasks. It can modify partial or overall video content and supports multi-image reference replacement, as well as replicating actions, effects, and camera movements. Wan2.7 - video editing |
| Image to video | 2026-04-03 | wan2.7-i2v | The Wan 2.7 image-to-video model supports multimodal input (text, image, audio, and video) for three main tasks: first-frame video generation, start-and-end-frame video generation, and video continuation. Wan-Image-to-Video 2.7 |
| Text-to-Video | 2026-04-03 | wan2.7-t2v | The Wan 2.7 text-to-video model adds new resolution options and custom aspect ratio settings for flexible adaptation to different creation scenarios and platform publishing requirements. Wan text-to-video |
| reference-to-video | 2026-04-03 | wan2.7-r2v | The Wan 2.7 reference-to-video model supports entity reference and voice customization. It also generates playbook-based videos directly from a single multi-panel storyboard. Wan - Reference-to-video |
| Reasoning model | 2026-04-02 | qwen3.6-plus, qwen3.6-plus-2026-04-02 | Qwen3.6-Plus features significantly enhanced code development capabilities (Agentic Coding, front-end programming, etc.) with a notably improved Vibe Coding experience. General reasoning capabilities are further strengthened. For multimodal tasks, object recognition, OCR, and object localization are significantly improved. Known issues from the Qwen3.5-Plus launch have also been fixed. The usage is the same as qwen3.5-plus. Overview |
| Image generation and editing | 2026-04-01 | wan2.7-image-pro, wan2.7-image | The Wan 2.7 image generation and editing model supports text-to-image, text-to-image group generation, image-to-image group generation, image editing, multi-image reference generation, and interactive editing. It excels in text rendering, entity consistency, and following complex instructions. The Pro series supports 4K outputs. The accelerated version balances quality and response speed. Wan - image generation and editing 2.7 |
| Omni-modal | 2026-03-30 | qwen3.5-omni-plus, qwen3.5-omni-plus-2026-03-15, qwen3.5-omni-flash, qwen3.5-omni-flash-2026-03-15 | The latest omni-modal model supports long video analysis, meeting minutes, caption output, content moderation, and audio-video interaction. It supports deep understanding and description generation for audio and video content, detects 113 languages, and generates audio in 36 languages. The model can process up to 3 hours of audio and 1 hour of video input. It also supports web search and instructions to control the volume, speed, and emotion of the audio output. Non-real-time (Qwen-Omni) |
| Omni-modal | 2026-03-30 | qwen3.5-omni-plus-realtime, qwen3.5-omni-plus-realtime-2026-03-15, qwen3.5-omni-flash-realtime, qwen3.5-omni-flash-realtime-2026-03-15 | The latest real-time multimodal model from Qwen. Compared to the previous generation Qwen3-Omni-Flash-Realtime, it features significantly improved intelligence on par with Qwen3.5-Plus. It natively supports web search (WebSearch), voice interruption, and voice control. It recognizes 113 languages and dialects, and generates speech in 36 languages and dialects. Real-time (Qwen-Omni-Realtime) |
| Reasoning model | 2026-03-20 | deepseek-v3.2 | DeepSeek-V3.2 is the official model that introduces DeepSeek Sparse Attention, a sparse attention mechanism. It is the first model from DeepSeek that integrates thinking with tool calling. The model supports tool calling in both thinking and non-thinking modes.DeepSeek |
| Image generation and editing | 2026-03-03 | qwen-image-2.0, qwen-image-2.0-2026-03-03, qwen-image-2.0-pro, qwen-image-2.0-pro-2026-03-03 | The Qwen Image 2.0 series supports both image generation and editing. The Pro series offers stronger text rendering, realistic quality, and semantic instruction following. The accelerated version balances quality and response speed. Qwen - text-to-image, Qwen-Image Edit |
| Speech recognition | 2026-03-03 | qwen3-asr-flash-2026-02-10 | A new snapshot model for Qwen audio file recognition provides better performance than qwen3-asr-flash-2025-09-08. Non-real-time speech recognition |
| Reasoning model | 2026-02-24 | qwen3.5-flash, qwen3.5-flash-2026-02-23, qwen3.5-122b-a10b, qwen3.5-27b, qwen3.5-35b-a3b | The latest Qwen3.5-Flash model and open-source model from Alibaba, supporting text, image, and video input with fast response speed. Its overall performance is close to qwen3.5-plus. Supports built-in Tool calling. Overview |
| Code model | 2026-02-20 | qwen3-coder-next | A new open-source code generation model in the Qwen3 series. It supports multi-turn tool interactions to enhance its understanding of repository-level code and improve its compatibility with AI programming tools. Code Capabilities (Qwen-Coder) |
| Reasoning model | 2026-02-16 | qwen3.5-plus, qwen3.5-plus-2026-02-15, qwen3.5-397b-a17b | The latest Qwen3.5-Plus and open-source models from Alibaba. They accept text, image, and video inputs and deliver outstanding performance across multiple tasks, such as language understanding, logical reasoning, code generation, agent tasks, image understanding, video understanding, and graphical user interface (GUI). They also support built-in tool calling. Overview |
| Speech recognition | 2026-02-13 | qwen3-asr-flash-realtime-2026-02-10 | A new snapshot model for Qwen real-time speech recognition performs better than qwen3-asr-flash-realtime-2025-10-27. Real-time speech recognition |
| Speech synthesis | 2026-02-10 | cosyvoice-v3-plus, cosyvoice-v3-flash | CosyVoice speech synthesis adds v3 models that support speech synthesis with system voices and cloned voices. Real-time speech synthesis - CosyVoice |
| Speech synthesis | 2026-02-10 | qwen3-tts-instruct-flash, qwen3-tts-instruct-flash-2026-01-26 | Qwen speech synthesis introduces an instruction-control model that supports precise control of synthesis output through natural language instructions. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-10 | qwen3-tts-vd-2026-01-26 | Qwen speech synthesis introduces a voice design model that supports creating customized voices through text descriptions. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-10 | qwen3-tts-vc-2026-01-22 | Qwen speech synthesis introduces a voice cloning model that quickly clones voices from real audio samples. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-04 | qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash-realtime-2026-01-22 | Qwen real-time speech synthesis adds the Instruct model, which supports precise control over synthesis results using natural language instructions. Qwen real-time speech synthesis |
| Reference-to-video | 2026-02-02 | wan2.6-r2v-flash | Generate multi-shot videos with automatic dubbing, using a character's appearance from a reference video or image. Wan2.7 - reference-to-video |
| Visual understanding | 2026-01-28 | qwen3-vl-flash-2026-01-22 | This is a new snapshot of the Qwen-VL model. It merges thinking and non-thinking modes to improve overall performance compared to the October 15, 2025 snapshot. The model achieves higher inference accuracy in scenarios such as general visual recognition, security, store patrols, inspections, and solving problems from photos. Image and video understanding |
| Speech recognition | 2026-01-28 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17 | The Qwen3-ASR-Flash-Filetrans series models now support word-level timestamps. By setting the new enable_words parameter, you can get millisecond-level word and character alignment information and experience more semantically accurate, fine-grained sentence segmentation. Non-real-time speech recognition |
| Reasoning model | 2026-01-27 | qwen3-max-2026-01-23 | Compared to the snapshot from September 23, 2025, this model merges thinking and non-thinking modes to significantly improve its overall performance. In thinking mode, the model integrates three tools: web search, web extractor, and code interpreter. Using external tools during its thought process, it achieves higher accuracy on complex problems. OpenAI-compatible - Responses |
| Image-to-video | 2026-01-18 | wan2.6-i2v-flash | Generates videos with or without audio. The two video types are billed independently based on their respective billing rules. It also has multi-shot narrative and audio processing capabilities. Wan - Image-to-video - Based on first frame |
| Image editing | 2026-01-18 | qwen-image-edit-max, qwen-image-edit-max-2026-01-16 | The Qwen Image Edit Max series provides more stable and versatile editing capabilities, enhances industrial design and geometric inference, and improves character consistency and editing precision. Image editing - Qwen |
| Speech synthesis | 2026-01-16 | qwen3-tts-vc-realtime-2026-01-15 | Qwen real-time speech synthesis adds a new snapshot model with further optimized quality. The synthesized voice is more natural and closer to the original compared to qwen3-tts-vc-realtime-2025-11-27. Real-time speech synthesis - Qwen, Real-time speech synthesis |
| Text-to-image | 2026-01-12 | qwen-image-plus-2026-01-09 | A new snapshot model for Qwen text-to-image generation, which is a distilled and accelerated version of qwen-image-max that rapidly generates high-quality images. Qwen-text-to-image |
| Image-to-video | 2026-01-08 | wan2.2-kf2v-flash | The model generates a seamless and smooth dynamic video from the first and last frame images based on a prompt. Image-to-video: First and last frames |
| Speech recognition | 2026-01-06 | qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Qwen3-ASR-Flash supports an OpenAI compatible mode. Recognize audio files using Qwen |
| Text-to-image | 2025-12-31 | qwen-image-max, qwen-image-max-2025-12-30 | The Qwen text-to-image model Max series offers enhanced realism and naturalness compared to the Plus series. It effectively reduces AI-generated artifacts and excels in aspects such as human figure textures, texture details, and text rendering. Qwen - text-to-image |
| Image editing | 2025-12-23 | qwen-image-edit-plus-2025-12-15 | The latest snapshot model for Qwen-Image-Editing improves character consistency, industrial design capabilities, and geometric inference compared to the previous version. It also optimizes the match in spatial layout, texture, and style between the edited and original images, resulting in more precise edits. Image editing - Qwen |
| Text-to-image | 2025-12-22 | z-image-turbo | A lightweight text-to-image model that quickly generates high-quality images. It supports bilingual rendering in Chinese and English, complex semantic understanding, multiple styles and themes, and flexibly adapts to various resolutions and aspect ratios. Z-Image |
| Visual understanding | 2025-12-19 | qwen3-vl-plus-2025-12-19 | The new Qwen-VL snapshot model features improved instruction-following capabilities and lower latency. Image and video understanding |
| Speech recognition | 2025-12-19 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17, qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Added support for speech recognition in 9 more languages, including Czech and Danish. Non-real-time speech recognition |
| Speech recognition | 2025-12-17 | qwen3-asr-flash-realtime, qwen3-asr-flash-realtime-2025-10-27 | Speech recognition now supports nine additional languages, including Czech and Danish. Real-time speech recognition |
| Speech recognition | 2025-12-17 | qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Supports audio with any sample rate and sound channel. Non-real-time speech recognition |
| Speech recognition | 2025-12-17 | fun-asr-mtl, fun-asr-mtl-2025-08-25 | Supports speech recognition in 31 languages, including Chinese, English, Japanese, and Korean. This feature is ideal for Southeast Asia scenarios. Audio file recognition - Fun-ASR/Paraformer |
| Voice design | 2025-12-16 | qwen-voice-design | Qwen has released a voice design model for generating customized voices from text descriptions. Use this model with the qwen3-tts-vd-realtime-2025-12-16 model to generate speech in 10 languages. Qwen voice design |
| Speech synthesis | 2025-12-16 | qwen3-tts-vd-realtime-2025-12-16 (snapshot) | Qwen real-time speech synthesis released a new snapshot model that enables low-latency, high-stability real-time synthesis using generated voice timbres. It supports multilingual output, automatically adjusts the tone based on the text, and optimizes synthesis performance for complex text. Real-time speech synthesis - QwenReal-time speech synthesis |
| Text-to-image | 2025-12-16 | wan2.6-t2i | A new sync interface is added. It supports selecting custom dimensions within the constraints of total pixel area and aspect ratio. Wan text-to-image V2 API reference |
| Image generation and editing | 2025-12-16 | wan2.6-image | Supports image editing and mixed image-text output. Wan - Image generation and editing 2.6 |
| Image-to-video based on the first frame | 2025-12-16 | wan2.6-i2v | Adds the multi-shot narrative feature, which supports audio using automatic dubbing or custom audio files. Wan - image-to-video - first frame (2.1-2.6) |
| reference-to-video | 2025-12-16 | wan2.6-r2v | Generates a multi-shot video based on a character's appearance and voice from a reference video. It supports automatic dubbing. Wan2.7 - reference-to-video |
| Text-to-video | 2025-12-16 | wan2.6-t2v | Adds the multi-shot narrative feature, which supports audio using automatic dubbing or custom audio files. Text-to-video |
| Speech recognition | 2025-12-12 | fun-asr, fun-asr-2025-11-07 | Feature updates for Fun-ASR audio file recognition: - Supports singing recognition and full song transcription. For details, see . |
| Omni-modal | 2025-12-04 | qwen3-omni-flash-2025-12-01 | The latest Qwen Omni snapshot model increases the number of supported timbres to 49 and features a significant upgrade to its instruction-following capabilities, enabling it to efficiently understand text, images, audio, and video. Non-real-time (Qwen-Omni) |
| Real-time multimodal | 2025-12-04 | qwen3-omni-flash-realtime-2025-12-01 | The latest snapshot model for the Qwen-Omni real-time version offers low-latency multimodal interaction. This version increases the number of supported timbres to 49 and significantly improves the model's instruction-following ability and interactive experience. Real-time (Qwen-Omni-Realtime) |
| Speech translation | 2025-12-04 | qwen3-livetranslate-flash, qwen3-livetranslate-flash-2025-12-01 | Qwen3-LiveTranslate-Flash is an audio and video translation model that translates between 18 languages, such as Chinese, English, Russian, and French. It leverages visual context to improve translation accuracy and provides both text and speech output. Audio and video file translation (Qwen) |
| Multilingual translation | 2025-12-02 | qwen-mt-lite | A basic text translation model from Qwen. It supports translation between 31 languages. It offers a faster response and lower cost than qwen-mt-flash, making it suitable for latency-sensitive scenarios. Machine translation (Qwen-MT) |
| Voice cloning | 2025-11-27 | qwen-voice-enrollment | Qwen released a voice cloning model. It generates a highly similar voice from just over 5 seconds of audio. When used with the qwen3-tts-vc-realtime-2025-11-27 model, it can create a high-fidelity clone of a person's voice and output it in real time across 11 languages. Qwen voice cloning |
| Speech synthesis | 2025-11-27 | qwen3-tts-vc-realtime-2025-11-27 (snapshot) | Qwen real-time speech synthesis released a new snapshot model that enables low-latency, high-stability real-time synthesis using generated voice timbres. It supports multilingual output, automatically adjusts the tone based on the text, and optimizes synthesis performance for complex text. Real-time speech synthesis - QwenReal-time speech synthesis |
| Speech synthesis | 2025-11-27 | qwen3-tts-flash-realtime-2025-11-27 (snapshot) | Qwen real-time speech synthesis released a new snapshot model. It features low latency and high stability. The model offers a richer selection of voices, and each voice supports multilingual output. It automatically adjusts the tone based on the text and enhances synthesis performance for complex text. Real-time speech synthesis |
| Speech synthesis | 2025-11-27 | qwen3-tts-flash-2025-11-27 (snapshot) | Qwen speech synthesis released a new snapshot model. It offers a richer selection of voices. Each voice supports multilingual output. The model automatically adjusts the tone based on the text and optimizes synthesis for complex text. Qwen text-to-speech (TTS) |
| Text extraction | 2025-11-21 | qwen-vl-ocr-2025-11-20 (snapshot) | This snapshot of the Qwen text extraction model is based on the Qwen3-VL architecture and significantly improves document parsing and text localization. Text extraction |
| Speech recognition | 2025-11-20 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17 (snapshot) | Qwen audio file recognition released a new model. It is designed for asynchronous transcription of audio files and supports recordings up to 12 hours long. Non-real-time speech recognition |
| Speech recognition | 2025-11-19 | fun-asr-2025-11-07 (snapshot) | A new snapshot model is available for Fun-ASR audio file recognition. This model optimizes far-field voice activity detection (VAD) to improve recognition accuracy and stability. In addition to Chinese and English, the model now supports multiple Chinese dialects and Japanese. Audio file recognition - Fun-ASR/Paraformer |
| Multilingual translation | 2025-11-11 | qwen-mt-flash | Compared to qwen-mt-turbo, this model supports streaming incremental output and offers improved overall performance. Machine translation (Qwen-MT) |
| Image-to-video | 2025-11-10 | wan2.2-animate-move | Transfers the actions and expressions of a character from a template video to a single static image to generate a video of the character in motion. Wan - Image-to-action |
| Image-to-video | 2025-11-10 | wan2.2-animate-mix | This model replaces the main character in a reference video with a character from an image. It preserves the original video's scene, lighting, and tone for seamless character replacement. Wan - Video Character Swap |
| Reasoning model | 2025-11-03 | qwen3-max-preview | The thinking mode of the qwen3-max-preview model features significantly improved overall inference capabilities. It performs especially well in agent programming, common-sense reasoning, and tasks related to math, science, and general purposes. Deep thinking |
| Image editing | 2025-10-31 | qwen-image-edit-plus, qwen-image-edit-plus-2025-10-30 | Built on qwen-image-edit, this model has optimized inference performance and system stability. It significantly reduces the response time for image generation and editing and supports returning multiple images in a single request. Image editing - Qwen |
| Real-time speech recognition | 2025-10-27 | qwen3-asr-flash-realtime, qwen3-asr-flash-realtime-2025-10-27 | The Qwen real-time speech recognition model features automatic language detection. It can detect 11 language types and provides accurate transcription in complex audio environments. Real-time speech recognition |
| Visual understanding | 2025-10-21 | qwen3-vl-32b-thinking, qwen3-vl-32b-instruct | A 32B dense model from the Qwen3-VL series. Its overall performance is second only to the Qwen3-VL-235B model. It excels in document recognition and understanding, spatial intelligence and object recognition, and visual 2D detection/spatial reasoning. This makes it suitable for complex perception tasks in general scenarios. Image and video understanding |
| Visual understanding | 2025-10-16 | qwen3-vl-flash, qwen3-vl-flash-2025-10-15 | A small-scale visual understanding model from the Qwen3 series. It effectively combines thinking and non-thinking modes. Compared to the open-source Qwen3-VL-30B-A3B, it delivers better performance and a faster response time. Image and video understanding |
| Visual understanding | 2025-10-14 | qwen3-vl-8b-thinking, qwen3-vl-8b-instruct | An 8B dense model from the Qwen3-VL series. It uses less GPU memory and performs multimodal understanding and inference. It supports ultra-long contexts such as long videos and documents, and visual 2D/3D positioning. It also features comprehensive spatial intelligence and object recognition. Image and video understanding |
| Visual understanding | 2025-10-03 | qwen3-vl-30b-a3b-thinking, qwen3-vl-30b-a3b-instruct | Based on the new-generation open-source Qwen3-VL model, this model responds quickly and features stronger multimodal understanding, inference, visual agent capabilities, and support for ultra-long contexts such as long videos and long documents. It also provides comprehensively upgraded spatial intelligence and object recognition to handle complex real-world tasks. Image and video understanding |
| Visual understanding | 2025-09-23 | qwen3-vl-plus, qwen3-vl-plus-2025-09-23, qwen3-vl-235b-a22b-thinking, qwen3-vl-235b-a22b-instruct | The Qwen3 visual understanding model effectively combines thinking and non-thinking modes, and its visual agent capabilities are among the best in the world. This version features comprehensive upgrades in visual encoding, spatial intelligence, and multimodal thinking. It also significantly improves visual perception and detection capabilities. Image and video understanding |
| Text-to-image | 2025-09-23 | qwen-image-plus | This model excels at rendering complex text, especially for Chinese and English. It creates complex layouts that mix images and text. It is also more cost-effective than qwen-image. Qwen - text-to-image |
| Code model | 2025-09-23 | qwen3-coder-plus-2025-09-23 | Compared to the previous version (snapshot from July 22), this model has improved performance on downstream tasks and greater robustness in tool calling. It also features enhanced code security. Code Capabilities (Qwen-Coder) |
| Reasoning model | 2025-09-11 | qwen-plus-2025-09-11 | This model is part of the Qwen3 series. Compared to qwen-plus-2025-07-28, it offers improved instruction following and generates more concise summaries in thinking mode. See deep thinking. In non-thinking mode, it has enhanced Chinese comprehension and logical reasoning. See Overview. |
| Reasoning model | 2025-09-11 | qwen3-next-80b-a3b-thinking, qwen3-next-80b-a3b-instruct | A next-generation open-source model based on Qwen3. The thinking modelhas improved instruction following and more concise summaries compared with qwen3-235b-a22b-thinking-2507. See Deep thinking.The instruct modelhas enhanced Chinese comprehension, logical reasoning, and text generation compared with qwen3-235b-a22b-instruct-2507. See Overview. |
| Text-to-text | 2025-09-05 | qwen3-max-preview | Qwen-Max is a preview model based on Qwen3. It offers a significant improvement in general capabilities over the Qwen 2.5 series. The model shows greatly enhanced performance in understanding Chinese and English text, following complex instructions, and handling subjective open-ended tasks. Its multilingual and tool calling capabilities are also stronger. The model is less prone to knowledge-based hallucinations. Qwen-Max |
| Image editing | 2025-08-19 | qwen-image-edit | The Qwen image editing model supports precise text editing in both Chinese and English, color adjustment, detail enhancement, style transfer, adding or removing objects, and changing positions and actions. It can perform complex image and text editing. Image editing - Qwen |
| Visual understanding | 2025-08-18 | qwen-vl-plus-2025-08-15 | A visual understanding model. It offers significant improvements in object detection, localization, and multilingual processing. Image and video understanding |
| Text-to-image | 2025-08-14 | qwen-image | The Qwen-Image model excels at rendering complex text, especially Chinese and English, to create intricate layouts that combine images and text.Text-to-image (Qwen-Image) |
| Visual understanding | 2025-08-13 | qwen-vl-max-2025-08-13 | A visual understanding model. It has comprehensively improved visual understanding metrics and significantly enhanced capabilities in math, reasoning, object recognition, and multilingual processing. Image and video understanding |
| Code model | 2025-08-05 | qwen3-coder-flash, qwen3-coder-flash-2025-07-28 | The fastest and most cost-effective model in the Qwen Coder series. Code Capabilities (Qwen-Coder). |
| Reasoning model | 2025-08-05 | qwen-flash, qwen-flash-2025-07-28 | The fastest and most cost-effective model in the Qwen series, suitable for simple jobs. |
| Reasoning model | 2025-07-30 | qwen-plus-2025-07-28 | A model from the Qwen3 series. Compared to the previous version, it increases the context length to 1,000,000. For more information about thinking mode, see Deep thinking. For more information about non-thinking mode, see Overview. |
| Reasoning model | 2025-07-30 | qwen3-30b-a3b-thinking-2507qwen3-30b-a3b-instruct-2507 | An upgrade of qwen3-30b-a3b. The thinking modelhas improved logical, general, knowledge, and creative capabilities. See Deep thinking.The instruct modelhas improved creative capabilities and model safety. See Overview. |
| Image-to-video | 2025-07-28 | wan2.2-i2v-plus | Compared to the 2.1 model, the new version significantly improves image detail and motion stability. The generation speed increases by up to 50%. First-frame video generation |
| Text-to-video | 2025-07-28 | wan2.2-t2v-plus | Compared to the 2.1 model, this new version has significantly improved image detail and motion stability. Generation speed is 50% faster. Text-to-video |
| Text-to-image | 2025-07-28 | wan2.2-t2i-flash, wan2.2-t2i-plus | This new version is a major upgrade to the 2.1 model. It is more creative, stable, and realistic. Images generate 50% faster. Text-to-image |
| Reasoning model | 2025-07-24 | qwen3-235b-a22b-thinking-2507, qwen3-235b-a22b-instruct-2507 | An upgrade of qwen3-235b-a22b. The thinking modelhas significantly improved logical, general, knowledge, and creative capabilities, suitable for challenging reasoning scenarios. See Deep thinking.The instruct modelhas improved creative capabilities and model safety. See Overview. |
| Code model | 2025-07-23 | qwen3-coder, qwen3-coder-plus-2025-07-22 | A code generation model based on Qwen3. It has strong Coding Agent capabilities and excels at tool calling and environment interaction. It combines excellent code capabilities with general-purpose abilities. Code Capabilities (Qwen-Coder) |
| Visual understanding | 2025-06-04 | qwen-vl-plus-2025-05-07 | A visual understanding model with significantly improved capabilities in mathematics, reasoning, and surveillance video content understanding. Image and video understanding |
| Text-to-image | 2025-05-22 | wan2.1-t2i-turbo, wan2.1-t2i-plus | Generates an image from a single sentence. The model generates images of any resolution and aspect ratio, up to 2 million pixels. It is available in two versions: turbo and plus. Text-to-image |
| Visual understanding | 2025-05-16 | qwen-vl-max-2025-04-08 | A visual understanding model. It has improved math and reasoning abilities. Its response style is tailored to human preferences, and its responses are significantly more detailed and clearly formatted. Image and video understanding |
| Visual understanding | 2025-05-16 | qwen-vl-plus-2025-01-25 | A visual understanding model belonging to the Qwen2.5-VL series. Compared to the previous version, this model expands the context to 128k and significantly enhances image and video understanding. |
| Video editing | 2025-05-19 | wan2.1-vace-plus | A general-purpose video editing model. The model supports multi-modal input, combining images, videos, and text prompts. It performs various tasks such as image-to-video (generating a video based on the entity or background of a reference image) and video repainting (extracting motion features from an input video to generate a video). General video editing |
| Reasoning model | 2025-04-28 | Qwen3 commercial modelsqwen-plus-2025-04-28, qwen-turbo-2025-04-28Qwen3 open-source models**qwen3-235b-a22b, qwen3-30b-a3b, qwen3-32b, qwen3-14b, qwen3-8b, qwen3-4b, qwen3-1.7b, qwen3-0.6b | Qwen3 models support both thinking and non-thinking modes. You can switch between the two modes by using the enable_thinking parameter. In addition, the capabilities of Qwen3 models have been greatly enhanced: - 1. Inference capability : In evaluations for math, code, and logical reasoning, the models significantly outperform QwQ and other non-reasoning models of similar size, reaching top-tier industry levels. - 2. Human preference alignment : Capabilities for creative writing, role-playing, multi-turn conversation, and instruction following are greatly improved. The general capabilities significantly exceed those of models of a similar size. - 3. Agent capability : The models achieve industry-leading performance in both reasoning and non-reasoning modes. They can accurately call external tools. - 4. Multilingual capability : The models support over 100 languages and dialects. Capabilities for multilingual translation, instruction understanding, and common-sense reasoning are significantly improved. - 5. Response format fixes : This version fixes response format issues from previous versions, such as abnormal Markdown, mid-sentence truncation, and incorrect boxed output. For more information about thinking mode, see Deep thinking. For more information about non-thinking mode, see Overview. |
| Text-to-video | 2025-04-21 | wan2.1-t2v-turbo, wan2.1-t2v-plus | - Generates a video from a single sentence. - It accurately follows instructions, supports large and complex motion, and simulates real-world physics. The generated videos feature rich artistic styles and cinematic quality. Wan text-to-video. |
| Image-to-video | 2025-04-21 | wan2.1-kf2v-plus, wan2.1-i2v-turbo, wan2.1-i2v-plus, | - Generates a smooth, dynamic video from start and end frame images based on a prompt. Image-to-video: Start and end frames - Takes an image as the first frame and generates a video based on prompts. For usage instructions, see first-frame-to-video.. |
| Visual reasoning | 2025-03-28 | qvq-max, qvq-max-latest, qvq-max-2025-03-25 | A visual reasoning model. It supports visual input and chain-of-thought output, demonstrating stronger capabilities in math, programming, visual analysis, creation, and general tasks. Visual reasoning |
| Omni-modal | 2025-03-26 | qwen2.5-omni-7b | A new multimodal model from Qwen for understanding and generation. It supports text, image, speech, and video input, and outputs text and audio. It provides two natural conversational voices. Non-real-time (Qwen-Omni). |
| Visual understanding | 2025-03-24 | qwen2.5-vl-32b-instruct | A visual understanding model. Its ability to solve math problems is close to the level of Qwen2.5VL-72B. The response style has been significantly adjusted to align with human preferences. The model now gives much more detailed and clearly formatted answers, especially for objective questions in math, logical reasoning, and knowledge Q\&A. Image and video understanding |
| Reasoning model | 2025-03-06 | qwq-plus | The QwQ reasoning model, trained on the Qwen2.5 model, greatly improves its reasoning ability through reinforcement learning. The model's core metrics for math and code (AIME 24/25, LiveCodeBench) and some general metrics (IFEval, LiveBench, etc.) reach the level of the full-performance DeepSeek-R1. Deep thinking |
| Visual understanding | 2025-01-27 | qwen2.5-vl-3b-instructqwen2.5-vl-7b-instructqwen2.5-vl-72b-instruct | - Compared to the Qwen2-VL model, this model has the following improvements: It has significantly improved capabilities in instruction following, mathematical calculations, code generation, and structured output (JSON output). - It supports unified parsing of visual content in images, such as text, charts, and layouts. The model can also accurately locate visual elements and represent them using detection frames and coordinate points. - It understands long video files (up to 10 minutes), provides event localization with second-level precision, and recognizes the sequence and speed of events. - See Image and video understanding. |
| Text-to-text | 2025-01-27 | qwen-max-2025-01-25qwen2.5-14b-instruct-1mqwen2.5-7b-instruct-1m | - The qwen-max-2025-01-25 model (also known as Qwen2.5-Max): The best-performing model in the Qwen series. It has significant improvements in code writing and comprehension, logical reasoning, and multilingual capabilities. The response style is greatly adjusted to align with human preferences. The detail and format clarity of its responses are significantly improved, with targeted enhancements in content creation, JSON format adherence, and role assumption. See Overview of text generation models. - The qwen2.5-14b-instruct-1m and qwen2.5-7b-instruct-1m models: Compared to the qwen2.5-14b-instruct and qwen2.5-7b-instruct models, the context length is increased to 1,000,000. See Overview of text generation models. |
| Text-to-text | 2025-01-17 | qwen-plus-2025-01-12 | - Compared to the qwen-plus-2024-12-20 model, this model has enhanced overall capabilities in both Chinese and English, with significant improvements in common sense and reading comprehension. Its ability to switch naturally between different languages, dialects, and styles is significantly improved, and Chinese instruction following is significantly enhanced. See Qwen. |
| Multilingual translation | 2024-12-25 | qwen-mt-plusqwen-mt-turbo | - The Qwen-MT model is a machine translation large language model optimized based on the Qwen model. It excels at translation between Chinese and English, Chinese and minor languages, and English and minor languages. The minor languages include 26 languages such as Japanese, Korean, French, Spanish, German, Portuguese (Brazil), Thai, Indonesian, Vietnamese, and Arabic. In addition to multilingual translation, the model supports features such as terminology intervention, domain prompting, and translation memory to improve translation quality in complex scenarios. See Machine translation (Qwen-MT). |
| Visual understanding | 2024-12-18 | qwen2-vl-72b-instruct | - Achieved state-of-the-art results on multiple visual understanding benchmarks, significantly enhancing the processing of multimodal tasks. See Image and video understanding for instructions. |
US
If you select the US deployment scope, model inference compute resources are restricted to the United States. Static data is stored in your selected region. Supported region: US (Virginia).
| Type | Time | Model | Description |
|---|---|---|---|
| Multilingual translation | 2026-04-10 | qwen-mt-lite-us | A basic-tier Qwen text translation model supporting mutual translation across 31 languages. Compared with qwen-mt-flash, it provides faster responses at lower cost, suitable for latency-sensitive scenarios. Machine translation (Qwen-MT) |
| Visual understanding | 2026-03-14 | qwen3-vl-flash-2026-01-22-us | A new snapshot model for Qwen-VL that effectively integrates thinking mode and non-thinking mode. Compared with the snapshot version from October 15, 2025, it significantly improves overall model performance, achieving higher inference accuracy in scenarios such as general visual recognition, security, store inspection, patrol inspection, and photo-based problem solving. Image and video understanding |
| Image-to-video based on the first frame | 2026-01-04 | wan2.6-i2v-us | Added multi-shot narrative capability. Supports audio capabilities, including automatic dubbing or custom audio file input. Wan - image-to-video - first frame (2.1-2.6) |
| Text-to-video | 2026-01-04 | wan2.6-t2v-us | Added multi-shot narrative capability. Supports audio capabilities, including automatic dubbing or custom audio file input. Wan - text-to-video |
| Speech recognition | 2026-01-04 | qwen3-asr-flash-us, qwen3-asr-flash-2025-09-08-us | Supports audio with any sample rate and number of channels. Non-real-time speech recognition |
| Visual understanding | 2026-01-04 | qwen3-vl-flash-us, qwen3-vl-flash-2025-10-15-us | A small-size visual understanding model in the Qwen3 series that effectively integrates thinking mode and non-thinking mode. Compared with the open-source Qwen3-VL-30B-A3B, it delivers better results with faster response speed. Image and video understanding |
| Reasoning model | 2026-01-04 | qwen-plus-2025-12-01-us | This is a model from the Qwen3 series. Compared to qwen-plus-2025-07-28, it has improved instruction-following capabilities and provides more concise summary responses in thinking mode. See Deep thinking. In non-thinking mode, its Chinese understanding and logical reasoning abilities are enhanced. See Overview. |
| Reasoning model | 2026-01-04 | qwen-plus-us | Offers a balance of capabilities. Its inference performance, cost, and speed are between those of Qwen-Max and Qwen-Flash, making it suitable for moderately complex tasks. |
| Text-to-text | 2026-01-04 | qwen-flash-us, qwen-flash-2025-07-28-us | The fastest and most cost-effective model in the Qwen series, suitable for simple jobs. |
Chinese mainland
If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).
| Type | Time | Model | Description |
|---|---|---|---|
| Music generation | 2026-05-06 | fun-music-v1 | The Fun Music model accepts open-ended song creation requests or lyrics as input to generate full songs with male or female vocals in Chinese or English. Songs are catchy and emotionally rich, combining human inspiration with large model capabilities. Music generation |
| Text-to-video | 2026-04-27 | happyhorse-1.0-t2v | HappyHorse text-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - text-to-video |
| Image-to-video | 2026-04-27 | happyhorse-1.0-i2v | HappyHorse image-to-video model. Supports audio video generation with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - image-to-video-first frame |
| Reference-to-video | 2026-04-27 | happyhorse-1.0-r2v | HappyHorse reference-to-video model. Supports multiple reference images as input to generate audio videos with 3-15 seconds duration at 720P/1080P resolution. HappyHorse - reference-to-video |
| Video editing | 2026-04-27 | happyhorse-1.0-video-edit | HappyHorse video editing model. Supports video editing and processing. HappyHorse - video editing |
| Text-to-video | 2026-04-26 | wan2.7-t2v-2026-04-25 | A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same as wan2.7-t2v. Wan - text-to-video |
| Image-to-video | 2026-04-26 | wan2.7-i2v-2026-04-25 | A snapshot version of the Wan 2.7 image-to-video model. The model capabilities are the same as wan2.7-i2v. Wan2.7 - image-to-video |
| Reasoning model | 2026-04-24 | deepseek-v4-pro, deepseek-v4-flash | DeepSeek-V4 series models hosted by Alibaba Cloud (Beijing region). deepseek-v4-pro is a large-scale MoE model with strong general reasoning capabilities. deepseek-v4-flash is a lightweight and cost-effective model optimized for speed.DeepSeek |
| Image generation | 2026-04-23 | qwen-image-2.0-pro-2026-04-22 | A model in the Qwen-Image-2.0 series that unifies image generation and editing. Compared with the March 3 snapshot, this model delivers a significant leap in visual quality --- especially texture detail, lighting, and materials. It supports multilingual in-image text generation and delivers more balanced artistic style performance. |
| Reasoning model | 2026-04-23 | qwen3.5-plus-2026-04-20 | A new snapshot of the Qwen3.5 native vision-language Plus model. Compared with the February 15 snapshot, agentic coding capabilities are significantly improved and inference speed is notably faster. Knowledge, reasoning, and long-context capabilities remain at a high level, making it well suited for agentic coding, production workflows, and high-throughput scenarios. |
| Reasoning model | 2026-04-23 | qwen3.6-27b | A 27B native vision-language dense model in the Qwen3.6 series. Compared with Qwen3.5-27B, agentic coding capabilities are significantly improved, and STEM and reasoning abilities are further enhanced. On the vision side, spatial intelligence and object localization & detection are notably stronger, while video understanding, document OCR, and visual Agent capabilities are steadily improved. |
| Text-to-text & image understanding | 2026-04-21 | kimi-k2.6 | The latest and most intelligent model from Moonshot AI, with stronger and more stable long-range code writing capabilities, and significantly improved instruction following and self-correction abilities. Supports text, image, and video input, thinking and non-thinking modes, as well as dialogue and Agent tasks. Kimi |
| Reasoning model | 2026-04-20 | qwen3.6-max-preview | The largest proprietary model in the Qwen3.6 series, with further improved coding capabilities and more efficient agent execution. Supports text-only input, thinking mode (enabled by default), explicit caching, and function calling. Overview ** Image and video inputs are not supported. |
| Reasoning model | 2026-04-16 | qwen3.6-flash, qwen3.6-flash-2026-04-16, qwen3.6-35b-a3b | A native vision-language Flash model in the Qwen3.6 series, delivering significant overall performance improvements over Qwen3.5-Flash. Agentic coding capabilities are substantially enhanced (far surpassing the predecessor on multiple code agent benchmarks), along with stronger math and code reasoning. On the vision side, spatial intelligence is notably improved, with object localization and detection being particularly outstanding. Overview |
| Speech recognition | 2026-04-08 | fun-asr, fun-asr-2025-11-07 | Fun-ASR real-time speech recognition capability upgrades: - Dialect support: Covers the seven major Chinese dialects (Mandarin, Wu, Xiang, Gan, Hakka, Min, Yue) and 20+ regional accents - Ancient poetry optimization: Improves ancient poetry recognition accuracy for cultural education and audiobook scenarios - Text optimization: Enhances punctuation prediction and text normalization. Numbers, dates, and amounts automatically convert to standard formats - Multi-language expansion: Supports 30 languages, including Chinese, English, Japanese, Korean, French, German, and Spanish For more information, see Audio file recognition - Fun-ASR/Paraformer. |
| Reasoning model | 2026-04-14 | glm-5.1 | The Zhipu GLM-5.1 model is designed for long-horizon tasks and supports a 200K context window with a maximum output of 128K tokens. With strong logical reasoning, long-text comprehension, and code generation capabilities, it delivers excellent results across multiple benchmarks and is well suited for intelligent interaction, enterprise applications, and development assistance. GLM-Alibaba Cloud |
| Text-to-text | 2026-04-14 | qwen-flash-character | A multilingual role-playing model in the Qwen series. Designed for lifelike character personas. Strictly follows the assigned persona, proactively drives conversations, listens and empathizes. Deeply recreates personalized characters. Role-playing (Qwen-Character) |
| Video editing | 2026-04-03 | wan2.7-videoedit | The Wan 2.7 video editing model supports instruction-based editing and video migration. It can modify parts of a video or the entire frame, and supports multi-image reference replacement and the replication of actions, effects, and camera movements. Wan2.7 - video editing |
| Image-to-video | 2026-04-03 | wan2.7-i2v | The Wan 2.7 image-to-video model supports multimodal input, including text, images, audio, and video. It performs three main tasks: generating video from a first frame, generating video from first and last frames, and video continuation. Wan2.7 - image-to-video |
| Text-to-video | 2026-04-03 | wan2.7-t2v | The Wan 2.7 text-to-video model adds new resolution tiers and custom aspect ratio settings. This allows for flexible adaptation to different creative scenarios and platform publishing needs. Wan - text-to-video |
| Reference-to-video | 2026-04-03 | wan2.7-r2v | The Wan 2.7 reference-to-video model supports subject referencing and voice customization. It can generate a scripted video directly from a single, multi-panel storyboard image. Wan2.7 - reference-to-video |
| Reasoning model | 2026-04-02 | qwen3.6-plus, qwen3.6-plus-2026-04-02 | Qwen3.6-Plus with significant upgrades in coding capabilities (Agentic Coding, frontend programming, etc.) and notably improved Vibe Coding experience. General-purpose reasoning is further enhanced. Multimodal capabilities including universal recognition, OCR, and object localization are significantly improved. Known issues from the Qwen3.5-Plus release are also fixed. Usage is the same as qwen3.5-plus. Overview |
| Image generation and editing | 2026-04-01 | wan2.7-image-pro, wan2.7-image | The Wan 2.7 image generation and editing model supports text-to-image, text-to-image-set, image-to-image-set, image editing, multi-image reference generation, and interactive editing, with improved performance in text rendering, subject consistency, and complex instruction following. The Pro series supports 4K output; the accelerated version balances quality and response speed. Wan - image generation and editing 2.7 |
| Omni-modal | 2026-03-30 | qwen3.5-omni-plus, qwen3.5-omni-plus-2026-03-15, qwen3.5-omni-flash, qwen3.5-omni-flash-2026-03-15 | The latest omni-modal model. It supports long video analysis, meeting minutes, subtitle output, security audits, and audio-video interaction. It also supports deep understanding and description generation for audio and video content, recognizes 113 languages, and generates audio in 36 languages. It can process up to 3 hours of audio and 1 hour of video input, and supports web search and instructions to control the volume, speed, and emotion of the output audio. Non-real-time (Qwen-Omni) |
| Omni-modal | 2026-03-30 | qwen3.5-omni-plus-realtime, qwen3.5-omni-plus-realtime-2026-03-15, qwen3.5-omni-flash-realtime, qwen3.5-omni-flash-realtime-2026-03-15 | The latest real-time multimodal model from Qwen. Compared to the previous generation Qwen3-Omni-Flash-Realtime, it features significantly improved intelligence on par with Qwen3.5-Plus. It natively supports web search (WebSearch), voice interruption, and voice control. It recognizes 113 languages and dialects, and generates speech in 36 languages and dialects. Real-time (Qwen-Omni-Realtime) |
| Reasoning model | 2026-03-11 | MiniMax-M2.5 | A new model from MiniMax, featuring fast response speeds and excelling in tasks such as coding and office work. Usage |
| Speech recognition | 2026-03-05 | fun-asr-realtime-2026-02-28 | Fun-ASR real-time speech recognition adds a new snapshot model, which performs better than fun-asr-realtime-2025-11-07. Real-time speech recognition - Fun-ASR/Paraformer |
| Image generation and editing | 2026-03-03 | qwen-image-2.0, qwen-image-2.0-2026-03-03, qwen-image-2.0-pro, qwen-image-2.0-pro-2026-03-03 | The Qwen Image 2.0 series supports both image generation and editing. The Pro series offers stronger text rendering, realistic quality, and semantic instruction following. The accelerated version balances quality and response speed. Qwen - text-to-image, Qwen-Image Edit |
| Speech recognition | 2026-03-03 | qwen3-asr-flash-2026-02-10 | Qwen audio file recognition adds a new snapshot model, which performs better than qwen3-asr-flash-2025-09-08. Non-real-time speech recognition |
| Speech synthesis | 2026-03-02 | cosyvoice-v3.5-plus, cosyvoice-v3.5-flash | The CosyVoice3.5 model is released. It focuses on voice cloning and design, and supports instruction-based control of speech synthesis. Real-time speech synthesis - CosyVoice |
| Reasoning model | 2026-02-24 | qwen3.5-flash, qwen3.5-flash-2026-02-23, qwen3.5-122b-a10b, qwen3.5-27b, qwen3.5-35b-a3b | The latest Qwen3.5-Flash model and open-source model from Alibaba, supporting text, image, and video input with fast response speed. Its overall performance is close to qwen3.5-plus. Supports built-in Text generation overview. Overview |
| Code model | 2026-02-20 | qwen3-coder-next | A new-generation open source code generation model in the Qwen3 series. It supports multi-turn tool interactions, and has an enhanced understanding of repository-level code and better adaptability to AI programming tools. Code capabilities (Qwen-Coder) |
| Reasoning model | 2026-02-18 | glm-5 | Zhipu's latest model, designed for programming and agent scenarios, excels at complex systems engineering and long-horizon agent tasks. GLM-Alibaba Cloud |
| Reasoning model | 2026-02-16 | qwen3.5-plus, qwen3.5-plus-2026-02-15, qwen3.5-397b-a17b | The latest Qwen3.5-Plus and open source models from Alibaba. They support text, image, and video inputs and deliver outstanding performance across multiple tasks, such as language understanding, logical reasoning, code generation, agent tasks, image understanding, video understanding, and graphical user interface (GUI) interaction. They also support built-in tool calling. Overview |
| Speech recognition | 2026-02-13 | qwen3-asr-flash-realtime-2026-02-10 | A new snapshot model for Qwen real-time speech recognition is added, which performs better than qwen3-asr-flash-realtime-2025-10-27. Real-time speech recognition |
| Speech recognition | 2026-02-12 | fun-asr-flash-8k-realtime, fun-asr-flash-8k-realtime-2026-01-28 | A new lightweight ASR model based on the Fun-ASR architecture, optimized for 8 kHz scenarios and suitable for cost-sensitive customers. Real-time speech recognition - Fun-ASR/Paraformer |
| Speech synthesis | 2026-02-10 | qwen3-tts-instruct-flash, qwen3-tts-instruct-flash-2026-01-26 | Qwen speech synthesis introduces an Instruct model, which lets you precisely control the synthesis effect using natural language instructions. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-10 | qwen3-tts-vd-2026-01-26 | Qwen speech synthesis introduces a voice design model, which lets you create custom voices using text descriptions. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-10 | qwen3-tts-vc-2026-01-22 | Qwen speech synthesis introduces a voice cloning model, which lets you quickly clone voices from real audio samples. Non-real-time speech synthesis |
| Speech synthesis | 2026-02-04 | qwen3-tts-instruct-flash-realtime, qwen3-tts-instruct-flash-realtime-2026-01-22 | Qwen real-time speech synthesis adds an Instruct model, which lets you precisely control the synthesis effect using natural language instructions. Real-time speech synthesis |
| Reference-to-video | 2026-02-02 | wan2.6-r2v-flash | Generates multi-shot videos with automatic dubbing, using a character's appearance from a reference video or image. Wan - Reference-to-Video |
| Text generation and visual understanding | 2026-01-30 | kimi-k2.5 | A visual understanding model from Moonshot AI that excels at general intelligence tasks such as code generation and visual understanding. It supports image, video, and text inputs, and can perform dialogue and agent tasks. Kimi |
| Speech recognition | 2026-01-28 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17 | The Qwen3-ASR-Flash-Filetrans series models now support word-level timestamps. By setting the new enable_words parameter, you can get millisecond-level word and character alignment information and experience more semantically accurate, fine-grained sentence segmentation. Non-real-time speech recognition |
| Reasoning model | 2026-01-27 | qwen3-max-2026-01-23 | Compared to the snapshot version from September 23, 2025, this model effectively merges thinking and non-thinking modes, significantly enhancing its overall performance. In thinking mode, the model is integrated with three tools: web search, web information extraction, and code interpreter. Using external tools during its thought process, it achieves higher accuracy on complex problems. OpenAI-compatible - Responses |
| Visual understanding | 2026-01-23 | qwen3-vl-flash-2026-01-22 | A new snapshot of Qwen-VL. Compared to the snapshot from October 15, 2025, it effectively merges thinking and non-thinking modes, significantly enhancing the model's overall performance. It achieves higher inference accuracy in business scenarios such as general visual recognition, security monitoring, store and patrol inspections, and photo-based problem solving. Image and video understanding |
| Image-to-video | 2026-01-17 | wan2.6-i2v-flash | Supports the generation of videos with and without sound. The two types of videos are billed independently according to their respective billing rules. It also provides multi-shot narrative and audio processing capabilities. Wan - Image-to-video - based on first frame |
| Image editing | 2026-01-17 | qwen-image-edit-max, qwen-image-edit-max-2026-01-16 | The Qwen-Image-Edit-Max series offers more stable and richer editing capabilities, enhanced industrial design and geometric inference capabilities, and improved character consistency and editing precision. Image editing - Qwen |
| Speech synthesis | 2026-01-16 | qwen3-tts-vc-realtime-2026-01-15 | Qwen real-time speech synthesis adds a new snapshot model with further optimized quality. The synthesized voice is more natural and closer to the original compared to qwen3-tts-vc-realtime-2025-11-27. Real-time speech synthesis - Qwen, Real-time speech synthesis |
| Text-to-image | 2026-01-12 | qwen-image-plus-2026-01-09 | This new snapshot model for Qwen text-to-image is a distilled and accelerated version of qwen-image-max that rapidly generates high-quality images. Qwen - text-to-image |
| Reasoning model | 2026-01-12 | deepseek-v3.2 | The deepseek-v3.2 model supports implicit and explicit caching to improve response speed and reduce usage costs without affecting response quality. Context cache |
| Image-to-video | 2026-01-08 | wan2.2-kf2v-flash | Based on a prompt and the input start and end frames, the model can generate a smooth, dynamic video. First-and-last-frame-to-video |
| Speech recognition | 2026-01-06 | qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Qwen3-ASR-Flash supports OpenAI compatible mode. Non-real-time speech recognition |
| Speech synthesis | 2026-01-05 | cosyvoice-v3-flash | CosyVoice speech synthesis adds 24 new voice timbres (for details, see Voice list): - Dialects: Long Jiayi, Long Laotie - Overseas marketing: loongkyong, loongtomoka - Poetry reading: Long Fei - Voice assistant: Long Xiaochun, Long Xiaoxia, YUMI - Social companionship: Long Cheng, Long Ze, Long Zhe, Long Yan, Long Xing, Long Tian, Long Wan, Long Qiang, Long Feifei, Long Hao - Audiobook: Long Sanshu, Long Yuan, Long Yue, Long Xiu, Long Nan - News report: Long Shu |
| Text-to-image | 2025-12-31 | qwen-image-max, qwen-image-max-2025-12-30 | The Max series of the Qwen image generation model enhances image realism and naturalness compared to the Plus series, effectively reduces AI-generated artifacts, and excels in areas such as character textures, texture details, and text rendering. Qwen - text-to-image |
| Image editing | 2025-12-23 | qwen-image-edit-plus-2025-12-15 | The latest snapshot model for Qwen image editing enhances character consistency, industrial design capabilities, and geometric inference compared to the previous version. It also optimizes the alignment of spatial layout, texture, and style between the edited and original images, resulting in more precise edits. Image editing - Qwen |
| Text-to-image | 2025-12-19 | z-image-turbo | A lightweight text-to-image model that quickly generates high-quality images. It supports bilingual rendering in Chinese and English, complex semantic understanding, multiple styles and themes, and flexible adaptation to various resolutions and aspect ratios. Text-to-image Z-Image |
| Visual understanding | 2025-12-19 | qwen3-vl-plus-2025-12-19 | The new Qwen-VL snapshot model features improved instruction-following capabilities and lower latency. Image and video understanding |
| Speech recognition | 2025-12-19 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17, qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Speech recognition now supports 9 more languages, including Czech and Danish. Non-real-time speech recognition |
| Speech recognition | 2025-12-17 | qwen3-asr-flash-realtime, qwen3-asr-flash-realtime-2025-10-27 | Speech recognition is now supported in nine additional languages, including Czech and Danish. Real-time speech recognition |
| Speech recognition | 2025-12-17 | qwen3-asr-flash, qwen3-asr-flash-2025-09-08 | Supports audio with any sample rate and sound channel. Non-real-time speech recognition |
| Speech recognition | 2025-12-17 | fun-asr-mtl, fun-asr-mtl-2025-08-25 | Supports speech recognition for 31 languages, including Chinese, English, Japanese, and Korean. Ideal for scenarios in Southeast Asia. Audio file recognition - Fun-ASR/Paraformer |
| Speech synthesis | 2025-12-16 | qwen3-tts-vd-realtime-2025-12-16 (snapshot) | Qwen real-time speech synthesis released a new snapshot model that enables low-latency, high-stability real-time synthesis using generated voice timbres. It supports multilingual output, automatically adjusts the tone based on the text, and optimizes synthesis performance for complex text. Real-time speech synthesis - QwenReal-time speech synthesis |
| Text-to-image | 2025-12-16 | wan2.6-t2i | A new sync interface is added. It supports selecting custom dimensions within the constraints of total pixel area and aspect ratio. Wan - Text-to-image V2 |
| Image generation and editing | 2025-12-16 | wan2.6-image | Supports image editing and mixed image-text output. Wan - Image Generation and Editing 2.6 |
| Image-to-video - based on first frame | 2025-12-16 | wan2.6-i2v | Adds the multi-shot narrative feature and audio support, allowing you to use automatic dubbing or upload a custom audio file. Wan - Image-to-video - based on first frame |
| Reference-to-video | 2025-12-16 | wan2.6-r2v | Generates a multi-shot video based on a character's appearance and voice from a reference video. It also supports automatic dubbing. Wan - Reference-to-Video |
| Text-to-video | 2025-12-16 | wan2.6-t2v | Adds the multi-shot narrative feature, which supports audio using automatic dubbing or custom audio files. Text-to-Video |
| Speech recognition | 2025-12-12 | fun-asr, fun-asr-2025-11-07 | Feature updates for Fun-ASR audio file recognition: - Supports singing recognition and full song transcription. For details, see . |
| Speech synthesis | 2025-12-11 | cosyvoice-v3-flash, cosyvoice-v3-plus | - The cosyvoice-v3-flash model now includes five new system voices: longanrou_v3, longyingjing_v3, longyingling_v3, longanling_v3, and longhan_v3. All these voices support the timestamp and SSML features. See Voice list. - The voice cloning feature for the cosyvoice-v3-flash and cosyvoice-v3-plus models has been enhanced. The feature now supports timestamps and SSML, and provides improved prosody. To try the enhancement, create a new voice. See CosyVoice voice cloning/design API. |
| Omni-modal | 2025-12-04 | qwen3-omni-flash-2025-12-01 | The latest Qwen Omni snapshot model increases the number of supported timbres to 49 and features a significant upgrade to its instruction-following capabilities, enabling it to efficiently understand text, images, audio, and video. Non-real-time (Qwen-Omni) |
| Real-time multimodal | 2025-12-04 | qwen3-omni-flash-realtime-2025-12-01 | The latest snapshot model for the Qwen Omni real-time version offers low-latency multimodal interaction. The number of supported timbres is increased to 49, and the model's instruction-following ability and interactive experience are significantly upgraded. Real-time (Qwen-Omni-Realtime) |
| Speech translation | 2025-12-04 | qwen3-livetranslate-flash, qwen3-livetranslate-flash-2025-12-01 | Qwen3-LiveTranslate-Flash is an audio and video translation model that translates between 18 languages, such as Chinese, English, Russian, and French. It uses visual context to improve translation accuracy and provides both text and speech output. Audio and video file translation - Qwen |
| Reasoning model | 2025-12-04 | deepseek-v3.2 | DeepSeek-V3.2 is the official model that introduces DeepSeek Sparse Attention, a sparse attention mechanism. It is also the first DeepSeek model to integrate thinking with tool usage, and it supports tool calling in both thinking and non-thinking modes.DeepSeek |
| Multilingual translation | 2025-12-02 | qwen-mt-lite | The Qwen basic text translation model supports translation between 31 languages. Compared to qwen-mt-flash, this model provides faster responses at a lower cost, making it suitable for latency-sensitive scenarios. Machine translation (Qwen-MT) |
| Voice cloning | 2025-11-27 | qwen-voice-enrollment | Qwen released a voice cloning model. It can quickly generate a highly similar voice from an audio clip of 5 seconds or more. When combined with the qwen3-tts-vc-realtime-2025-11-27 model, it can clone a person's voice with high fidelity and output it in real time across 11 languages. Voice cloning (Qwen) |
| Speech synthesis | 2025-11-27 | qwen3-tts-vc-realtime-2025-11-27 (snapshot) | Qwen real-time speech synthesis released a new snapshot model that enables low-latency, high-stability real-time synthesis using generated voice timbres. It supports multilingual output, automatically adjusts the tone based on the text, and optimizes synthesis performance for complex text. Real-time speech synthesis - QwenReal-time speech synthesis |
| Speech synthesis | 2025-11-27 | qwen3-tts-flash-realtime-2025-11-27 (snapshot) | Qwen real-time speech synthesis released a new snapshot model. It offers low latency and high stability. It provides more voice options, and a single voice can support multilingual output. The model automatically adjusts tone based on the text and improves synthesis for complex text. Real-time speech synthesis |
| Speech synthesis | 2025-11-27 | qwen3-tts-flash-2025-11-27 (snapshot) | Qwen speech synthesis released a new snapshot model. It offers more voice options, and a single voice supports multilingual output. The model automatically adapts its tone to the text and provides optimized synthesis for complex text. Non-real-time speech synthesis |
| Text extraction | 2025-11-21 | qwen-vl-ocr-2025-11-20 (snapshot) | This snapshot of the Qwen text extraction model is based on the Qwen3-VL architecture and significantly improves document parsing and text localization. Text extraction |
| Speech recognition | 2025-11-20 | qwen3-asr-flash-filetrans, qwen3-asr-flash-filetrans-2025-11-17 (snapshot) | Qwen Audio File Recognition released a new model. This model is designed for the asynchronous transcription of audio files and supports recordings up to 12 hours long. Non-real-time speech recognition |
| Speech synthesis | 2025-11-19 | cosyvoice-v3-flash | This version improves pronunciation accuracy and voice similarity compared to previous versions. It also supports more languages, including German, Spanish, French, Italian, and Russian. Real-time speech synthesis - CosyVoice |
| Reasoning model | 2025-11-11 | kimi-k2-thinking | This is a thinking model from Moonshot AI. It has general agent and reasoning capabilities, excels at deep reasoning, and can solve complex problems through multi-step tool calling. Kimi |
| Multilingual translation | 2025-11-10 | qwen-mt-flash | Compared to qwen-mt-turbo, this model supports streaming incremental output and has improved overall performance. Machine translation (Qwen-MT) |
| Image-to-video | 2025-11-04 | wan2.2-animate-mix | Replaces the main character in a reference video with the character from an input image, while preserving the original video's scene, lighting, and tone for a seamless character replacement. Wan - Video character replacement |
| Reasoning model | 2025-11-03 | qwen3-max-preview | The thinking mode of the qwen3-max-preview model provides significant improvements in overall reasoning capabilities, especially delivering superior performance in agent programming, common-sense reasoning, math, science, and general tasks. Deep thinking |
| Image-to-video | 2025-11-03 | wan2.2-animate-move | Transfers the actions and expressions of a character from a template video to a single static character image to generate a video of the character in motion. Wan - Image-to-action |
| Image editing | 2025-10-31 | qwen-image-edit-plus, qwen-image-edit-plus-2025-10-30 | Based on qwen-image-edit, this model offers optimized inference performance and system stability. It greatly reduces the response time for image generation and editing, and supports returning multiple images in a single request. Image editing - Qwen |
| Real-time speech recognition | 2025-10-27 | qwen3-asr-flash-realtime, qwen3-asr-flash-realtime-2025-10-27 | The Qwen real-time speech recognition model features automatic language detection. It can recognize 11 language types and transcribe audio accurately in complex environments. Real-time speech recognition |
| Visual understanding | 2025-10-21 | qwen3-vl-32b-thinking, qwen3-vl-32b-instruct | A 32B dense model from the Qwen3-VL series. Its overall performance is second only to the Qwen3-VL-235B model. It excels in document recognition and understanding, spatial intelligence, object recognition, 2D visual detection, and spatial reasoning. It is suitable for complex perception tasks in general scenarios. Image and video understanding |
| Visual understanding | 2025-10-16 | qwen3-vl-flash, qwen3-vl-flash-2025-10-15 | A small-sized visual understanding model from the Qwen3 series. It effectively combines thinking and non-thinking modes. Compared to the open source Qwen3-VL-30B-A3B, it performs better and responds faster. Image and video understanding |
| Visual understanding | 2025-10-14 | qwen3-vl-8b-thinking, qwen3-vl-8b-instruct | An 8B dense open source model from the Qwen3-VL series, available in both thinking and non-thinking versions. It uses less GPU memory and can perform multimodal understanding and reasoning. It supports ultra-long contexts such as long videos and long documents, 2D/3D visual positioning, and comprehensive spatial intelligence and object recognition. Image and video understanding |
| Visual understanding | 2025-10-03 | qwen3-vl-30b-a3b-thinking, qwen3-vl-30b-a3b-instruct | Based on the new-generation, open source Qwen3-VL model, it is available in both thinking and non-thinking versions. It has a fast response speed and stronger capabilities for multimodal understanding, reasoning, and visual agent tasks. It also supports ultra-long contexts such as long videos and documents. Its spatial intelligence and object recognition capabilities are fully upgraded to handle complex real-world tasks. Image and video understanding |
| Reasoning model | 2025-09-30 | deepseek-v3.2-exp | A hybrid reasoning model that supports both thinking and non-thinking modes. It introduces a sparse attention mechanism to improve training and inference efficiency for long texts, priced lower than deepseek-v3.1. For details, see . |
| Text-to-image | 2025-09-23 | qwen-image-plus | This model excels at rendering complex text, especially Chinese and English. It can create complex mixed-media layouts of images and text. It is more cost-effective than qwen-image. Text-to-image (Qwen-Image) |
| Visual understanding | 2025-09-23 | qwen3-vl-plus, qwen3-vl-plus-2025-09-23, qwen3-vl-235b-a22b-thinking, qwen3-vl-235b-a22b-instruct | The Qwen3 series of visual understanding models effectively combines thinking and non-thinking modes, and their visual agent capabilities are world-class. This version features comprehensive upgrades in visual encoding, spatial intelligence, and multimodal thinking. Their visual perception and recognition capabilities are significantly improved. Image and video understanding |
| Code model | 2025-09-23 | qwen3-coder-plus-2025-09-23 | Compared to the previous version (July 22 snapshot), this version has improved robustness for downstream tasks and tool calling, and enhanced code security. Code capabilities (Qwen-Coder) |
| Reasoning model | 2025-09-11 | qwen-plus-2025-09-11 | This model is part of the Qwen3 series. Compared to qwen-plus-2025-07-28, it follows instructions better and provides more concise summaries in thinking mode. See Deep thinking. In non-thinking mode, it has enhanced Chinese language understanding and logical reasoning. See Overview. |
| Reasoning model | 2025-09-11 | qwen3-next-80b-a3b-thinking, qwen3-next-80b-a3b-instruct | A next-generation open-source model based on Qwen3. The thinking modelhas improved instruction following and more concise summaries compared with qwen3-235b-a22b-thinking-2507. See Deep thinking.The instruct modelhas enhanced Chinese comprehension, logical reasoning, and text generation compared with qwen3-235b-a22b-instruct-2507. See Overview. |
| Text generation | 2025-09-05 | qwen3-max-preview | The Qwen-Max model (preview), based on Qwen3, offers significant improvements in general capabilities compared to the Qwen 2.5 series. It has significantly enhanced abilities in Chinese and English text understanding, complex instruction following, subjective open-ended tasks, multilingual tasks, and tool calling. The model also has fewer knowledge hallucinations. |
| Text, image, video, voice, etc. | 2025-08-05 | Models | This is the first release in the Beijing region. |
China (Hong Kong)
If you select the Hong Kong (China) deployment scope, model inference compute resources are restricted to Hong Kong (China). Static data is stored in your selected region. Supported region: China (Hong Kong).
| Type** | Time | Model | Description |
|---|---|---|---|
| Reasoning model | 2026-03-17 | qwen3-max, qwen3-max-2026-01-23 | Compared to the snapshot from September 23, 2025, this model better integrates thinking and non-thinking modes for significantly improved overall performance. In thinking mode, the model uses three tools: Web search, web page information extraction, and a code interpreter. Using these external tools helps the model achieve higher accuracy on complex problems. OpenAI compatible-Responses |
| Reasoning model | 2026-03-17 | qwen-plus, qwen-plus-2025-12-01 | This model is part of the Qwen3 series and is an improvement over qwen-plus-2025-07-28. In thinking mode, it follows instructions better and provides more concise summaries. See Deep thinking. In non-thinking mode, its Chinese language understanding and logical reasoning are enhanced. See Overview. |
| Reasoning model | 2026-03-17 | qwen3.5-flash, qwen3.5-flash-2026-02-23 | The latest Qwen3.5-Flash model and open-source model from Alibaba, supporting text, image, and video input with fast response speed. Supports built-in Tool calling. Overview |
| Visual understanding | 2026-03-17 | qwen3-vl-plus, qwen3-vl-plus-2025-12-19 | A visual understanding model in the Qwen3 series that effectively integrates thinking mode and non-thinking mode, achieving world-leading visual agent capabilities. This version includes comprehensive upgrades in visual encoding, spatial perception, and multimodal reasoning. Visual perception and recognition capabilities are significantly improved, with stronger instruction-following ability and lower latency. Image and video understanding |
EU
If you select the European Union deployment scope, model inference compute resources are restricted to the European Union. Static data is stored in your selected region. TSupported region: Germany (Frankfurt).
| Type | Time | Model | Feature description |
|---|---|---|---|
| Reasoning model | 2026-03-20 | qwen3-max, qwen3-max-2026-01-23 | Compared with the snapshot version from September 23, 2025, this version effectively integrates thinking mode and non-thinking mode, significantly improving overall model performance. In thinking mode, the model integrates three tools: web search, web page information extraction, and code interpreter, achieving higher accuracy on complex problems by incorporating external tools into the reasoning process. OpenAI-compatible - Responses |
| Reasoning model | 2026-03-20 | qwen-plus, qwen-plus-2025-12-01 | This is a Qwen3 series model. Compared to qwen-plus-2025-07-28, it offers better instruction-following in thinking mode and provides more concise summary responses. See Deep thinking. In non-thinking mode, its Chinese language understanding and logical reasoning are enhanced. See Overview. |
| Reasoning model | 2026-03-20 | qwen3.5-flash, qwen3.5-flash-2026-02-23 | The latest Qwen3.5-Flash model and open-source model from Alibaba, supporting text, image, and video input with fast response speed. Supports built-in Tool calling. Overview |
| Visual understanding | 2026-03-20 | qwen3-vl-plus, qwen3-vl-flash, qwen3-vl-flash-2025-10-15 | A visual understanding model in the Qwen3 series that effectively integrates thinking mode and non-thinking mode, achieving world-leading visual agent capabilities. This version includes comprehensive upgrades in visual encoding, spatial perception, and multimodal reasoning. Visual perception and recognition capabilities are significantly improved, with stronger instruction-following ability and lower latency. Image and video understanding |
| Code capability | 2026-03-20 | qwen3-coder-next | The Qwen3 series' next-generation open-source code generation model supports multi-turn tool interaction, improves the understanding of repository-level code, and enhances adaptability to AI programming tools. Code Capabilities (Qwen-Coder) |