Appearance
Speech synthesis (Qwen-TTS)
Request parameters and response fields for the non-realtime speech synthesis (Qwen-TTS) API.
For usage instructions, see Non-real-time speech synthesis .
## Request body
## Non-streaming output
## Python
** The SpeechSynthesizer interface in the DashScope Python SDK is now unified under MultiModalConversation. Its usage and parameters remain fully consistent.
python
import os
import dashscope
# This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1
dashscope.base_http_api_url = 'https://dashscope-intl.aliyuncs.com/api/v1'
text = "Let me recommend a T-shirt to everyone. This one is really super nice. The color is very elegant, and it's also a perfect item to match. Everyone can buy it without hesitation. It's truly beautiful and very forgiving on the figure. No matter what body type you have, it will look great. I recommend everyone to place an order."
# SpeechSynthesizer interface usage: dashscope.audio.qwen_tts.SpeechSynthesizer.call(...)
response = dashscope.MultiModalConversation.call(
# To use the instruction control feature, replace the model with qwen3-tts-instruct-flash
model="qwen3-tts-flash",
# The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not configured, replace the following line with your Model Studio API key: api_key="sk-xxx"
api_key=os.getenv("DASHSCOPE_API_KEY"),
text=text,
voice="Cherry"
# To use the instruction control feature, uncomment the following line and replace the model with qwen3-tts-instruct-flash
# instructions='Fast speech rate, with a clear rising intonation, suitable for introducing fashion products.',
# optimize_instructions=True
)
print(response)## Java
java
// Install the latest version of the DashScope SDK
import com.alibaba.dashscope.aigc.multimodalconversation.AudioParameters;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.JsonUtils;
import com.alibaba.dashscope.utils.Constants;
public class Main {
// To use the instruction control feature, replace MODEL with qwen3-tts-instruct-flash
private static final String MODEL = "qwen3-tts-flash";
public static void call() throws ApiException, NoApiKeyException, UploadFileException {
MultiModalConversation conv = new MultiModalConversation();
MultiModalConversationParam param = MultiModalConversationParam.builder()
.model(MODEL)
// The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
// If the environment variable is not configured, replace the following line with your Model Studio API key: apiKey("sk-xxx")
.apiKey(System.getenv("DASHSCOPE_API_KEY"))
.text("Today is a wonderful day to build something people love!")
.voice(AudioParameters.Voice.CHERRY)
.languageType("English")
// To use the instruction control feature, uncomment the following lines and replace MODEL with qwen3-tts-instruct-flash
// .parameter("instructions","Fast speech rate, with a clear rising intonation, suitable for introducing fashion products.")
// .parameter("optimize_instructions",true)
.build();
MultiModalConversationResult result = conv.call(param);
System.out.println(JsonUtils.toJson(result));
}
public static void main(String\[\] args) {
// This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1
Constants.baseHttpApiUrl = "https://dashscope-intl.aliyuncs.com/api/v1";
try {
call();
} catch (ApiException \| NoApiKeyException \| UploadFileException e) {
System.out.println(e.getMessage());
}
System.exit(0);
}
}## curl
curl
# ======= IMPORTANT NOTE =======
# This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
# The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not configured, replace $DASHSCOPE_API_KEY with your Model Studio API key: sk-xxx.
# === DELETE THIS COMMENT WHEN EXECUTING ===
curl -X POST 'https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \\
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \\
-H 'Content-Type: application/json' \\
-d '{
"model": "qwen3-tts-flash",
"input": {
"text": "Let me recommend a T-shirt to everyone. This one is really super nice. The color is very elegant, and it's also a perfect item to match. Everyone can buy it without hesitation. It's truly beautiful and very forgiving on the figure. No matter what body type you have, it will look great. I recommend everyone to place an order.",
"voice": "Cherry",
"language_type": "Chinese"
}
}'## Streaming output
## Python
The SpeechSynthesizer interface in the DashScope Python SDK is now unified under MultiModalConversation. To switch to the new interface, simply replace the name --- all other parameters are fully compatible.
python
# DashScope SDK version 1.24.5 or later required
import os
import dashscope
# This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1
dashscope.base_http_api_url = 'https://dashscope-intl.aliyuncs.com/api/v1'
text = "Let me recommend a T-shirt to everyone. This one is really super nice. The color is very elegant, and it's also a perfect item to match. Everyone can buy it without hesitation. It's truly beautiful and very forgiving on the figure. No matter what body type you have, it will look great. I recommend everyone to place an order."
# SpeechSynthesizer interface usage: dashscope.audio.qwen_tts.SpeechSynthesizer.call(...)
response = dashscope.MultiModalConversation.call(
# To use the instruction control feature, replace the model with qwen3-tts-instruct-flash
model="qwen3-tts-flash",
# The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not configured, replace the following line with your Model Studio API key: api_key="sk-xxx"
api_key=os.getenv("DASHSCOPE_API_KEY"),
text=text,
voice="Cherry",
# To use the instruction control feature, uncomment the following lines and replace the model with qwen3-tts-instruct-flash
# instructions='Fast speech rate, with a clear rising intonation, suitable for introducing fashion products.',
# optimize_instructions=True,
stream=True
)
for chunk in response:
print(chunk)## Java
java
// DashScope SDK version 2.19.0 or later required
import com.alibaba.dashscope.aigc.multimodalconversation.AudioParameters;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversation;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationParam;
import com.alibaba.dashscope.aigc.multimodalconversation.MultiModalConversationResult;
import com.alibaba.dashscope.exception.ApiException;
import com.alibaba.dashscope.exception.NoApiKeyException;
import com.alibaba.dashscope.exception.UploadFileException;
import com.alibaba.dashscope.utils.JsonUtils;
import com.alibaba.dashscope.utils.Constants;
import io.reactivex.Flowable;
public class Main {
// To use the instruction control feature, replace MODEL with qwen3-tts-instruct-flash
private static final String MODEL = "qwen3-tts-flash";
public static void streamCall() throws ApiException, NoApiKeyException, UploadFileException {
MultiModalConversation conv = new MultiModalConversation();
MultiModalConversationParam param = MultiModalConversationParam.builder()
.model(MODEL)
// The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
// If the environment variable is not configured, replace the following line with your Model Studio API key: apiKey("sk-xxx")
.apiKey(System.getenv("DASHSCOPE_API_KEY"))
.text("Today is a wonderful day to build something people love!")
.voice(AudioParameters.Voice.CHERRY)
.languageType("English")
// To use the instruction control feature, uncomment the following lines and replace MODEL with qwen3-tts-instruct-flash
// .parameter("instructions","Fast speech rate, with a clear rising intonation, suitable for introducing fashion products.")
// .parameter("optimize_instructions",true)
.build();
Flowable<MultiModalConversationResult> result = conv.streamCall(param);
result.blockingForEach(r -\> {System.out.println(JsonUtils.toJson(r));
});
}
public static void main(String\[\] args) {
// This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1
Constants.baseHttpApiUrl = "https://dashscope-intl.aliyuncs.com/api/v1";
try {
streamCall();
} catch (ApiException \| NoApiKeyException \| UploadFileException e) {
System.out.println(e.getMessage());
}
System.exit(0);
}
}## curl
curl
# ======= IMPORTANT NOTE =======
# This is the URL for the Singapore region. If using a model in the Beijing region, replace the URL with: https://dashscope.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation
# The API keys for Singapore and Beijing regions are different. Get your API Key: https://www.alibabacloud.com/help/en/model-studio/get-api-key
# If the environment variable is not configured, replace $DASHSCOPE_API_KEY with your Model Studio API key: sk-xxx.
# === DELETE THIS COMMENT WHEN EXECUTING ===
curl -X POST 'https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/multimodal-generation/generation' \\
-H "Authorization: Bearer $DASHSCOPE_API_KEY" \\
-H 'Content-Type: application/json' \\
-H 'X-DashScope-SSE: enable' \\
-d '{
"model": "qwen3-tts-flash",
"input": {
"text": "Let me recommend a T-shirt to you. This one is truly stunning. Its color highlights your elegance and makes it an ideal match for any outfit. You can buy it without hesitation - it looks great on everyone. It flatters all body types. Whether you're tall, short, slim, or curvy, this T-shirt suits you perfectly. We highly recommend ordering it.",
"voice": "Cherry",
"language_type": "Chinese"
}
}'To play Base64-encoded audio in real time, see Speech synthesis - Qwen.
<b>model **`*string*`* ***(required)**The model name. For details, see Supported models. **input **`*object*`* ***(required)**Input parameters.
Properties
text *string* (required) The text to synthesize. Supports multilingual mixed input. Maximum input length: 512 tokens (Qwen-TTS model) or 600 characters (other models).
voice *string* (required) The voice to use. See Built-in voices.
language_type *string* (optional) The language of the synthesized audio. Defaults to Auto.
Auto: Use when the input contains multiple languages or the language cannot be determined. The model automatically matches pronunciation for each language segment, though accuracy is not guaranteed.- Specific language: Use for single-language text. Specifying the language significantly improves synthesis quality and typically produces better results than
Auto. Valid values: <li>Chinese EnglishGermanItalianPortugueseSpanishJapaneseKoreanFrenchRussian</li>
**instructions **string* *(optional) The instructions for speech synthesis. See Instruction-based control. Default: None. Maximum length: 1,600 tokens. Supported languages: Chinese and English only. Scope: This feature applies only to the Qwen3-TTS-Instruct-Flash-Realtime series models.
**optimize_instructions **boolean* *(optional)
When enabled, semantically optimizes the instructions to improve the naturalness and expressiveness of the synthesized speech. Default: false. Behavior: When set to true, the system semantically rewrites the instructions to generate directives better suited for speech synthesis. Use this parameter when precise control over speech delivery is needed. Depends on the instructions parameter. Has no effect if instructions is empty. Scope: This feature applies only to the Qwen3-TTS-Instruct-Flash series models.
## Response object (streaming and non-streaming formats are identical)
## Qwen3-TTS-Flash
json
{
"status_code": 200,
"request_id": "5c63c65c-cad8-4bf4-959d-xxxxxxxxxxxx",
"code": "",
"message": "",
"output": {
"text": null,
"finish_reason": "stop",
"choices": null,
"audio": {
"data": "",
"url": "http://dashscope-result-bj.oss-cn-beijing.aliyuncs.com/1d/ab/20251218/d2033070/39b6d8f2-c0db-4daa-9073-5d27bfb66b78.wav?Expires=1766113409\&OSSAccessKeyId=LTAI5xxxxxxxxxxxx\&Signature=NOrqxxxxxxxxxxxx%3D",
"id": "audio_5c63c65c-cad8-4bf4-959d-xxxxxxxxxxxx",
"expires_at": 1766113409
}
},
"usage": {
"input_tokens": 0,
"output_tokens": 0,
"characters": 195
}
}## Qwen-TTS
json
{
"status_code": 200,
"request_id": "f4e8139b-3203-4887-92cb-xxxxxxxxxxxx",
"code": "",
"message": "",
"output": {
"text": null,
"finish_reason": "stop",
"choices": null,
"audio": {
"data": "",
"url": "http://dashscope-result-wlcb.oss-cn-wulanchabu.aliyuncs.com/1d/50/20251218/e6c1b9cc/9acec74e-e317-4dbd-9e76-745c47bcbf2d.wav?Expires=1766116806\&OSSAccessKeyId=LTAxxxxxxxxx\&Signature=afYZxxxxxxxxx%2FAX9bk%3D",
"id": "audio_f4e8139b-3203-4887-92cb-xxxxxxxxxxxx",
"expires_at": 1766116806
}
},
"usage": {
"input_tokens": 76,
"output_tokens": 1045,
"characters": 0,
"input_tokens_details": {
"text_tokens": 76
},
"output_tokens_details": {
"audio_tokens": 1045,
"text_tokens": 0
},
"total_tokens": 1121
}
} **status_code** `*integer*`The HTTP status code as defined in RFC 9110. Common values:
• 200: Request succeeded. • 400: Invalid request parameters. • 401: Unauthorized. • 404: Resource not found. • 500: Internal server error. request_id *string*A unique identifier for this request. Use it for troubleshooting. code *string*The error code returned on failure. See Error messages. message *string*The error message returned on failure. See Error messages. **output ***object*The model output. Properties
**text ***string* Always null. Ignore this field.
**choices ***string* Always null. Ignore this field.
**finish_reason ***string* The generation status:
null--- Generation is in progress.stop--- Generation finished normally, or a stop condition was met.
audio *object* The audio output from the model. Properties
url *string* The URL of the complete audio file. Valid for 24 hours.
data *string* Base64-encoded audio data in streaming output.
id *string* A unique identifier for the audio.
expires_at *integer* The URL expiration time as a Unix timestamp.
usage *object*Token or character usage for this request. Qwen-TTS returns token usage; Qwen3-TTS-Flash returns character usage. Properties
input_tokens_details *object* Token usage details for the input text. Returned only by the Qwen-TTS model. Propertiestext_tokens *integer* The number of tokens consumed by the input text.
total_tokens *integer* The total number of tokens consumed by this request. Returned only by the Qwen-TTS model.
output_tokens *integer* The number of tokens consumed by the output audio. For the Qwen3-TTS-Flash model, this field is always 0.
input_tokens *integer* The number of tokens consumed by the input text. For the Qwen3-TTS-Flash model, this field is always 0.
output_tokens_details *object* Token usage details for the output. Returned only by the Qwen-TTS model. Properties
audio_tokens *integer* The number of tokens consumed by the output audio.
text_tokens *integer* The number of tokens consumed by the output text. Currently always 0.
characters *integer* The number of characters in the input text. Returned only by the Qwen3-TTS-Flash model.
request_id *string*A unique identifier for this request. Use it for troubleshooting.