Skip to content

Voice cloning Python SDK reference

CosyVoice voice cloning is accessible through the DashScope Python SDK. User guide: Voice cloning.

Service endpoint

The SDK uses the China (Beijing) region endpoint by default. To switch to a different region, set dashscope.base_http_api_url before you initialize the client.

International

If you select the International deployment scope, model inference compute resources are dynamically scheduled worldwide, excluding the Chinese mainland. Static data is stored in your selected region. Supported region: Singapore.

https://dashscope-intl.aliyuncs.com/api/v1

Chinese mainland

If you select the Chinese mainland deployment scope, model inference compute resources are restricted to the Chinese mainland. Static data is stored in your selected region. Supported region: China (Beijing).

https://dashscope.aliyuncs.com/api/v1

Switch to the Singapore region:

HELPCODEESCAPE-python
import dashscope


dashscope.base_http_api_url = 'https://dashscope-intl.aliyuncs.com/api/v1'

Note :

  • API keys differ between regions. Use the API key that corresponds to the target region.

  • The region setting is global and affects all DashScope SDK API calls.

VoiceEnrollmentService class

Package path : dashscope.audio.tts_v2.VoiceEnrollmentService

Purpose: Manage the lifecycle of CosyVoice cloned voices, including creating, querying, updating, and deleting voices.

Constructor

HELPCODEESCAPE-python
VoiceEnrollmentService()

create_voice() --- Create a voice

Method signature:

HELPCODEESCAPE-python
def create_voice(self, target_model: str, prefix: str, url: str,
                 language_hints: List[str] = None,
                 max_prompt_audio_length: float = None,
                 enable_preprocess: bool = None) -> str

Parameters:

ParameterTypeRequiredDescription
target_modelstrYesThe text-to-speech (TTS) model that drives the cloned voice. It must match the model you specify when calling the TTS API; otherwise, synthesis fails.
prefixstrYesA prefix for the voice name. Only alphanumeric characters are allowed, with a maximum length of 10 characters. The resulting voice name follows this format: {target_model}-{prefix}-{unique_id}.
urlstrYesThe URL of the audio file for voice cloning. The URL must be publicly accessible.
language_hintsList[str]No** **Important ** Applies only to CosyVoice voice cloning (when model is voice-enrollment). Supported only by cosyvoice-v3.5-plus, v3.5-flash, v3-plus, and v3-flash. Helps the model identify the language of the sample audio to extract voice features more accurately and improve cloning quality. If the specified language doesn't match the actual audio language (for example, setting en when the audio is in Chinese), the system ignores this value and detects the language automatically. This parameter is an array, but the current version processes only the first element. Valid values vary by model: - cosyvoice-v3-plus: zh: Chinese - en: English - fr: French - de: German - ja: Japanese - ko: Korean - ru: Russian - cosyvoice-v3.5-plus, cosyvoice-v3.5-flash, cosyvoice-v3-flash: zh: Chinese - en: English - fr: French - de: German - ja: Japanese - ko: Korean - ru: Russian - pt: Portuguese - th: Thai - id: Indonesian - vi: Vietnamese Default: ["zh"].
max_prompt_audio_lengthfloatNo** **Important ** Applies only to CosyVoice voice cloning (when model is voice-enrollment). Supported only by cosyvoice-v3.5-plus, v3.5-flash, and v3-flash. The maximum duration (in seconds) of the reference audio after preprocessing. Valid values: [3.0, 30.0]. Longer durations produce better results. Default: 10.0.
enable_preprocessboolNo** **Important ** Applies only to CosyVoice voice cloning (when model is voice-enrollment). Supported only by cosyvoice-v3.5-plus, v3.5-flash, and v3-flash. Whether to enable audio preprocessing (noise reduction, audio enhancement, and volume normalization). Enable this for recordings with background noise. Disable it for recordings in quiet environments to preserve the original voice characteristics. Default: false.

Return value : str --- The voice ID (voice_id).

list_voice() --- List voices

Method signature:

HELPCODEESCAPE-python
def list_voice(self, prefix: str = None, page_index: int = 0, page_size: int = 10) -> list

Parameters:

ParameterTypeRequiredDescription
prefixstrNoFilter by voice name prefix.
page_indexintNoPage index. Default: 0.
page_sizeintNoNumber of entries per page. Default: 10.

Return value : list --- A list of voices.

query_voice() --- Query voice details

Method signature:

HELPCODEESCAPE-python
def query_voice(self, voice_id: str) -> dict

Parameters:

ParameterTypeRequiredDescription
voice_idstrYesThe voice ID to query.

Return value : dict --- Voice details.

update_voice() --- Update a voice

Method signature:

HELPCODEESCAPE-python
def update_voice(self, voice_id: str, url: str, language_hints: List[str] = None,
                 max_prompt_audio_length: float = None, enable_preprocess: bool = None) -> None

Parameters:

ParameterTypeRequiredDescription
voice_idstrYesThe voice ID to update.
urlstrYesThe new audio file URL.
language_hintsList[str]NoLanguage hints for the sample audio.
max_prompt_audio_lengthfloatNoMaximum duration of the reference audio.
enable_preprocessboolNoWhether to enable audio preprocessing.

delete_voice() --- Delete a voice

Method signature:

HELPCODEESCAPE-python
def delete_voice(self, voice_id: str) -> None

Parameters:

ParameterTypeRequiredDescription
voice_idstrYesThe voice ID to delete.

Code examples

The following examples show how to use the VoiceEnrollmentService class for common voice cloning operations.

Create a voice

HELPCODEESCAPE-python
from dashscope.audio.tts_v2 import VoiceEnrollmentService

service = VoiceEnrollmentService()

# Avoid frequent calls. Each call creates a new voice, and you can't create
# more after reaching the quota limit.
voice_id = service.create_voice(
    target_model='cosyvoice-v3-plus',
    prefix='myvoice',
    url='https://your-audio-file-url'
    # language_hints=['zh'],
    # max_prompt_audio_length=10.0,
    # enable_preprocess=False
)

print(f"Request ID: {service.get_last_request_id()}")
print(f"Voice ID: {voice_id}")

List voices

HELPCODEESCAPE-python
from dashscope.audio.tts_v2 import VoiceEnrollmentService

service = VoiceEnrollmentService()

# Filter by prefix, or set to None to list all voices
voices = service.list_voices(prefix='myvoice', page_index=0, page_size=10)

print(f"Request ID: {service.get_last_request_id()}")
print(f"Found voices: {voices}")

Query a specific voice

HELPCODEESCAPE-python
from dashscope.audio.tts_v2 import VoiceEnrollmentService

service = VoiceEnrollmentService()
voice_id = 'cosyvoice-v3-plus-myvoice-xxxxxxxx'

voice_details = service.query_voice(voice_id=voice_id)

print(f"Request ID: {service.get_last_request_id()}")
print(f"Voice Details: {voice_details}")

Update a voice

HELPCODEESCAPE-python
from dashscope.audio.tts_v2 import VoiceEnrollmentService

service = VoiceEnrollmentService()
service.update_voice(
    voice_id='cosyvoice-v3-plus-myvoice-xxxxxxxx',
    url='https://your-new-audio-file-url'
)
print(f"Update submitted. Request ID: {service.get_last_request_id()}")

Delete a voice

HELPCODEESCAPE-python
from dashscope.audio.tts_v2 import VoiceEnrollmentService

service = VoiceEnrollmentService()
service.delete_voice(voice_id='cosyvoice-v3-plus-myvoice-xxxxxxxx')
print(f"Deletion submitted. Request ID: {service.get_last_request_id()}")

Mirror of Alibaba Cloud Model Studio docs for reference and RAG. Not affiliated with Alibaba Cloud.