Skip to content

User Guide: Models

Browse documentation for user guide: models. Pages organized by topic:

Get Started

  • What is Alibaba Cloud Model Studio — Alibaba Cloud Model Studio is a one-stop model service platform
  • Make your first API call to Qwen — This guide walks you through making your first API call to a Qwen model on Alibaba Cloud Model Studio
  • Recommended models — Alibaba Cloud Model Studio offers Qwen and third-party models for text, image, audio, and video
  • Rate limits — Model Studio enforces rate limits to ensure fair use across all RAM users, workspaces, and API keys under one Alibaba Cloud account

Billing

Text Generation

  • Overview — Text generation models produce clear and coherent text based on input prompts
  • Multi-turn conversations — Qwen API is stateless — pass conversation history via a messages array each request
  • Streaming output — Stream Qwen model responses token by token using Server-Sent Events over persistent HTTP connections
  • Deep thinking — Deep thinking models reason before responding, which improves accuracy on complex tasks such as logical reasoning and math
  • Structured output — Structured output (JSON mode) makes the model return a valid JSON string that your code can parse directly, without extra text like json tha
  • Partial mode — For scenarios like code completion and text continuation, you can generate new content starting from an existing text fragment (prefix)
  • Context cache — Inference requests to a large model often include overlapping input, such as in a multi-turn conversation or a series of questions about the
  • Batch inference — Process large volumes of requests asynchronously at 50% the cost of real-time inference

Tool Calling

  • Function calling — Large language models may struggle with tasks that require real-time data or complex math
  • Web search — Large language models (LLMs) rely on training data that has a knowledge cutoff date
  • Web extractor — LLMs cannot directly access web page data
  • Code Interpreter — Enable the built-in Python Code Interpreter when calling a model
  • Text-to-image search — The text-to-image search tool enables a model to search the Internet for relevant images based on a text description
  • Search by image — The search by image tool enables the model to search the Internet for visually similar images based on an input image
  • Knowledge retrieval — Large Language Models (LLMs) cannot answer questions about private data
  • MCP — The Model Context Protocol (MCP) enables large language models to use external tools and data

Best Practices

  • Integrate tools — Token Plan (Team Edition) supports extending AI coding tool capabilities through built-in model tools, including web search, code interprete
  • Integrate multimodal generation models — Image generation models must be integrated through each tool's extension mechanism (Skill, Slash Command, or Agent)

Best Practices

Changelog

  • Model lifecycle and updates — The following tables list model releases
  • Announcements — Stay current with Alibaba Cloud Model Studio updates: Qwen3-235B-A22B price cuts, Qwen3 model family releases, batch inference API support,

Clients And Developer Tools

  • OpenClaw — OpenClaw is an open source personal AI assistant platform that lets you interact with AI through various messaging channels
  • Hermes Agent — Hermes Agent is a terminal AI coding tool
  • Claude Code — Claude Code is a command-line AI coding assistant developed by Anthropic
  • OpenCode — OpenCode is a terminal AI coding tool that connects to Alibaba Cloud Model Studio via Pay-as-you-go, Coding Plan, or Token Plan (Team Editio
  • Cursor — Cursor is an AI coding IDE
  • Codex — Codex is a terminal AI coding assistant developed by OpenAI
  • Qwen Code — Qwen Code is a terminal-based AI coding tool that can be connected to Alibaba Cloud Model Studio via pay-as-you-go, Coding Plan, or Token Pl
  • Cherry Studio — Cherry Studio is an open-source AI desktop client
  • Chatbox — Chatbox is a cross-platform AI client that connects to Model Studio using Token Plan (Team Edition), Coding Plan, or Pay-as-you-go
  • Cline — Cline is a VS Code extension for AI-assisted coding
  • Qoder — Qoder is an agentic coding platform for software development that supports a desktop IDE, CLI, and JetBrains plugin, and can connect to Alib
  • Lingma — Lingma is an Alibaba Cloud smart coding assistant that provides a standalone IDE and can connect to Alibaba Cloud Model Studio via Coding Pl
  • Kilo CLI — Kilo CLI is the command-line client for Kilo Code
  • Call image and video generation APIs with Postman and cURL — Use Postman or cURL to call image or video generation APIs in Model Studio
  • Dify — Dify is an open-source platform for building AI applications
  • More tools — In addition to the tools listed above, Alibaba Cloud Model Studio supports any third-party programming tool that is compatible with the Open

Coding Plan

  • Coding plan overview — Coding Plan gives you access to models including Qwen, GLM, Kimi, and MiniMax in popular AI coding tools
  • web search — Add a web search tool to Coding Plan so the model can retrieve real-time information
  • add visual understanding capabilities — Some models in Model Studio Coding Plan, such as qwen3
  • FAQ — Connection and configurationCommon errors and solutionsError messagePossible causeSolution400 InvalidParameter: Range of input length should

Deployment

  • Model deployment — Deploying a provides a dedicated inference service that supports performance requirements such as high concurrency and low latency

Embedding And Rerank

  • Embedding — Convert text, images, and videos into dense or sparse vectors using Model Studio embedding models
  • Rerank — Retrieval systems prioritize speed, which can sometimes compromise the precision of search results

Fine

Image Editing

Image Generation And Editing

  • Text-to-image — Use the text-to-image API to generate images from text descriptions

Inference

  • Music generation — Fun-Music generates full-length songs with male or female vocals in Chinese or English
  • Speech-to-speech — Select a model for use cases like conversational speech and speech translation

Model Usage Statistics And Performance Monitoring

  • Model usage — Understand how Model Studio measures AI consumption across tokens, image zhang, and video seconds — and how free-tier quotas work before HTT

Omni

Security And Compliance

  • Permission management — Alibaba Cloud Model Studio offers granular access control at the console and model level, helping organizations manage users across multiple
  • Security certifications and privacy — Compliance certificationsAlibaba Cloud Model Studio has achieved SOC 2 compliance with an unqualified opinion, validating its controls for S
  • Alibaba Cloud Model Studio - Training data summary — Understand how Qwen and Wan models are built on multimodal corpora spanning text, images, video, and audio—with quality filters, deduplicati

Specialized Models

Speech

Speech Synthesis

  • Real-time speech synthesis — Real-time speech synthesis converts text into natural speech over a WebSocket connection
  • Non-real-time speech synthesis — Non-real-time speech synthesis converts text to speech (TTS) through an HTTP API
  • Voice cloning — Voice Cloning creates a highly realistic custom voice from a 10- to 20-second audio sample, with no model training required
  • Voice Design — Voice Design lets you create custom voices from natural language descriptions alone, with no audio samples required
  • SSML — Use SSML (Speech Synthesis Markup Language) to fine-tune speech characteristics such as speed, pauses, and pronunciation

Support

  • FAQ — Frequently asked questions about Alibaba Cloud Model Studio
  • Related agreements — Understand the legal and service commitments governing Alibaba Cloud Model Studio — covering Terms of Service, Model Inference SLA uptime gu

Third

Token Plan (Team Edition)

  • Token Plan overview — Token Plan (Team Edition) is an AI model subscription from Alibaba Cloud Model Studio
  • Quick start — Subscribe to Token Plan (Team Edition) and get started in three steps: choose a plan, obtain your API key, and configure your AI tool
  • FAQ — Frequently asked questions about Token Plan (Team Edition), covering purchasing, usage, metering, and performance

Transmission Security

Video Editing

  • Video Editing 2.7 — The Wanxiang video editing model supports editing operations on input videos, such as adding, deleting, or modifying content, replacing back
  • General video editing — The Wan unified video editing model supports multimodal input (text, image, and video) and provides five core capabilities: multi-image refe

Video Generation And Editing

  • Text-to-video — Wan I2V supports multi-modal input (text, images, and audio) and generates videos up to 15 seconds long at 1080P resolution
  • Image-to-video 2.7 — Wan2
  • Image-to-video: first and last frames — The Wan image-to-video model generates smooth videos from a first-frame image, a last-frame image, and an optional text prompt
  • Reference-to-video — Wan-R2V accepts multimodal input (text, image, video, and audio) to generate performance videos

Visual Understanding

  • Image and video understanding — Visual understanding models can answer questions based on the images or videos that you provide
  • Text extraction (Qwen-OCR) — Automate document digitization with Qwen-OCR — extract structured text from scanned files, tables, and receipts via OpenAI-compatible and Da
  • Visual reasoning — Visual reasoning models output their thinking process before answering

Mirror of Alibaba Cloud Model Studio docs for reference and RAG. Not affiliated with Alibaba Cloud.