Use the Qwen-Audio-TTS/CosyVoice iOS SDK to turn text into high-quality, expressive speech in your iOS apps.
User guide: For model overviews and voice selection, see Text-to-speech models.
Architecture highlights:
Qwen-Audio-TTS/CosyVoice supports two invocation modes: one-shot input and streaming input.
One-shot input: best for short-text synthesis or when you need SSML markup.
Starts a streaming speech synthesis task and opens a connection to the server.
Method signature:
Parameters:
Return value:
Returns an error code.
ticket JSON example: The following example does not list every field. Add the others as your code requires:
ticket fields:
parameters JSON example: The following example does not list every field. Add the others as your code requires:
parameters fields:
Sends text to be synthesized. Use this method together with startStreamInputTts.
After you call startStreamInputTts, use this method to push text continuously.
After all text is sent, call stopStreamInputTts or asyncStopStreamInputTts to signal end of input.
Method signature:
Parameters:
Return value:
Returns an error code.
Synchronous method. Notifies the server that all text has been sent, then blocks until every audio chunk is synthesized and TTS_EVENT_SYNTHESIS_COMPLETE is received.
The block timeout is controlled by complete_waiting_ms.
Method signature:
Return value:
Returns an error code.
Asynchronous method. Notifies the server that all text has been sent and returns immediately. Synthesis continues in the background.
Use TTS_EVENT_SYNTHESIS_COMPLETE to detect when synthesis finishes.
Method signature:
Return value:
Returns an error code.
Immediately drops the connection to the server and terminates the current synthesis task. After this method is called, no further audio data callbacks fire.
Method signature:
Return value:
Returns an error code.
Synchronous one-shot synthesis method. Sends the text and blocks while all audio data is received, then returns after synthesis finishes. You do not need to call stopStreamInputTts afterward.
This method enables SSML by default. To disable SSML, set the
Parameters:
The
Return value:
Returns an error code.
This method sends all text for synthesis asynchronously. It returns immediately without waiting for audio data. You do not need to call stopStreamInputTts afterward.
This method enables SSML by default. To disable SSML, set the
Parameters:
The
Return value:
Returns an error code.
Callback protocol for Qwen-Audio-TTS/CosyVoice streaming speech synthesis. Implement this protocol to receive synthesis events, audio data, and logs.
Method signature:
Parameters:
The SDK fires this callback repeatedly during synthesis. Read the audio data from the callback.
Method signature:
Parameters:
This callback delivers detailed SDK-internal logs to help with troubleshooting and debugging.
Method signature:
Parameters:
Event-type enum for Qwen-Audio-TTS/CosyVoice streaming speech synthesis.
Example task-failed response:
SDK log-level enum that controls log output.
Purpose: Embed XML tags in the input text to precisely control pronunciation, speech rate, pauses, and other prosody details.
Limitations: SSML is supported only with one-shot input (the playStreamInputTts and asyncPlayStreamInputTts methods). It is not supported with streaming input (the sendStreamInputTts method).
How to use: When you call playStreamInputTts or asyncPlayStreamInputTts, the SDK enables SSML automatically. Pass text that contains SSML tags directly in the
Purpose: Make the model read common math formulas and expressions correctly.
How to use: Pass text that contains math expressions in LaTeX format directly in the
NeoNui
Architecture highlights:
- Singleton pattern: get the global instance through
[StreamInputTts get_instance]. - Callback-driven: receive events and audio data through the StreamInputTtsDelegate protocol.
- JSON configuration: pass parameters as JSON strings.
Call flows
Qwen-Audio-TTS/CosyVoice supports two invocation modes: one-shot input and streaming input.
One-shot input: best for short-text synthesis or when you need SSML markup.
- playStreamInputTts() or asyncPlayStreamInputTts() — Send a complete text and start synthesis. The former is synchronous and returns after synthesis finishes; the latter is asynchronous and returns immediately after starting synthesis.
- onStreamInputTtsDataCallback() — Receive audio data.
- TTS_EVENT_SYNTHESIS_COMPLETE — Synthesis complete.
- startStreamInputTts() — Initialize the SDK and set the callback delegate and connection parameters.
- sendStreamInputTts() — Continuously send text fragments to be synthesized.
- onStreamInputTtsDataCallback() — Receive audio data.
- stopStreamInputTts() or asyncStopStreamInputTts() — Send the end-of-synthesis request. The former is synchronous and returns after synthesis finishes; the latter is asynchronous and returns immediately after sending the request.
- TTS_EVENT_SYNTHESIS_COMPLETE — Synthesis complete.
startStreamInputTts
Starts a streaming speech synthesis task and opens a connection to the server.
Method signature:
| Parameter | Type | Description |
|---|---|---|
ticket | char* | JSON string that holds authentication, connection, and debugging settings. |
parameters | char* | JSON string that holds the speech synthesis effect settings. |
sessionId | char* | Client-specified session ID. If omitted, the server generates one. |
logLevel | NuiSdkLogLevel | Print level for the SDK's internal logs. |
saveLog | BOOL | Whether to save logs locally. When set to YES, you must specify a path with debug_path and can cap the file size with max_log_file_size. |
| Field | Type | Required | Description |
|---|---|---|---|
url | string | Yes | Service address: wss://maas.qwencloudapi.com/api-ws/v1/inference |
apikey | string | Yes | API key. To limit exposure if a long-lived key leaks, use a short-lived API key instead. |
device_id | string | Yes | Unique identifier for the end user. Set this to the in-app user ID or a client-generated device identifier. The ID is used for log tracing and troubleshooting. |
complete_waiting_ms | int | No | Timeout, in milliseconds, for waiting on the synthesis-complete event (TTS_EVENT_SYNTHESIS_COMPLETE) after you call stopStreamInputTts. Default: 10000. |
debug_path | string | No | Local path where log files are stored. This field takes effect only when saveLog is set to YES in startStreamInputTts, playStreamInputTts, or asyncPlayStreamInputTts. In that case, you must set this path; otherwise, an error is returned.At most two log files are kept locally. |
max_log_file_size | int | No | Maximum size of a log file, in bytes. This field takes effect only when saveLog is set to YES in startStreamInputTts, playStreamInputTts, or asyncPlayStreamInputTts.Default: 104857600 (100 × 1024 × 1024 bytes, that is, 100 MiB). |
log_track_level | int | No | Filter level for logs delivered through the log callback (onStreamInputTtsLogTrackCallback).Default: 2. Valid values: 0 (LOG_LEVEL_VERBOSE), 1 (LOG_LEVEL_DEBUG), 2 (LOG_LEVEL_INFO), 3 (LOG_LEVEL_WARNING), 4 (LOG_LEVEL_ERROR), 5 (LOG_LEVEL_NONE, disables the callback). Note: log_track_level and logLevel (set through startStreamInputTts, playStreamInputTts, or asyncPlayStreamInputTts) together determine which logs reach the callback. A log fires the callback only when its level is at or above both thresholds. For example, if log_track_level is 2 (INFO) and logLevel is 3 (WARNING), only logs at WARNING or higher (level >= 3) are delivered. |
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | The model name. |
voice | string | Yes | The voice used for speech synthesis.
|
format | string | No | The audio encoding format. Valid values:
|
enable_audio_decoder | BOOL | No | Whether to enable the SDK's internal decoder. Default: NO.This parameter takes effect only when the audio encoding format is opus or mp3. When enabled, the SDK decodes opus or mp3 audio data into PCM data before returning it. |
volume | int | No | The volume level. Default value: 50. Valid values: [0, 100]. |
sample_rate | int | No | The audio sample rate in Hz. Valid values: 8000, 16000, 22050 (default), 24000, 44100, 48000. |
rate | float | No | The speech rate. Default value: 1.0. Valid values: [0.5, 2.0]. |
pitch | float | No | The pitch. Default value: 1.0. Valid values: [0.5, 2.0]. |
bit_rate | int | No | The audio bit rate in kbps. When the audio format is mp3 or opus, use bit_rate to adjust the bit rate.Default value: 32. Valid values: [6, 510]. |
enable_ssml | boolean | No | Whether to enable SSML. Default: false.
For the SSML usage restrictions (supported models, voices, and APIs), see Limitations. |
word_timestamp_enabled | boolean | No | Specifies whether to enable word-level timestamps. Default value: false. Available only in streaming output mode. Supported voices: cloned voices of qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, and cosyvoice-v3-plus, and system voices marked as supported in Qwen-Audio-TTS voice list, CosyVoice voice list. Cloned voices of other models do not support this feature. Note: Timestamps are returned in all_response of onStreamInputTtsEventCallback. |
seed | int | No | A random seed for controlling variation in the synthesis output. When the model version, text, voice, and other parameters are unchanged, using the same seed produces identical results. Default value: 0. Valid values: [0, 65535]. |
language_hints | array[string] | No | Specifies the target language for speech synthesis to improve output quality. Note: This parameter is an array, but the current version only processes the first element. Pass a single value. This parameter specifies the target language for speech synthesis. It's unrelated to the language of the audio sample used in voice cloning. To set the source language for a cloning task, see the voice cloning API reference. When digit pronunciation, abbreviation expansion, symbol reading, or minority-language synthesis doesn't meet expectations, use this parameter. For example:
Valid values: zh (Chinese), en (English), fr (French), de (German), ja (Japanese), ko (Korean), ru (Russian), pt (Portuguese), th (Thai), id (Indonesian), vi (Vietnamese), es (Spanish), it (Italian), ms (Malaysian), fil (Filipino), ar (Arabic). |
instruction | string | No | Controls synthesis characteristics such as dialect, emotion, or speaking style. For usage details, see Instruction control. |
enable_aigc_tag | boolean | No | Specifies whether to embed an AIGC watermark in the generated audio. When set to true, the watermark is embedded in audio files of supported formats (wav/mp3/opus). Default value: false. Only qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, and cosyvoice-v3-plus support this feature. |
aigc_propagator | string | No | Sets the ContentPropagator field in the AIGC watermark, identifying the content propagator. Takes effect only when enable_aigc_tag is true.Default value: UID. Only qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, and cosyvoice-v3-plus support this feature. |
aigc_propagate_id | string | No | Sets the PropagateID field in the AIGC watermark, uniquely identifying a specific propagation action. Takes effect only when enable_aigc_tag is true.Default value: The request ID of the current speech synthesis request. Only qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash, cosyvoice-v3-flash, and cosyvoice-v3-plus support this feature. |
hot_fix | object | No | Text hot-fix configuration for customizing the pronunciation of specified words or replacing text before synthesis. Fields:
Example: {"hot_fix": {"pronunciation": [{"天气": "tian1 qi4"}], "replace": [{"今天": "金天"}]}} |
enable_markdown_filter | BOOL | No | Whether to filter Markdown markup from the input text before synthesis so that the markup is not spoken aloud. Note: Only cloned voices of cosyvoice-v3-flash support this feature. Default: NO.Valid values:
|
sendStreamInputTts
Sends text to be synthesized. Use this method together with startStreamInputTts.
After you call startStreamInputTts, use this method to push text continuously.
After all text is sent, call stopStreamInputTts or asyncStopStreamInputTts to signal end of input.
Method signature:
| Parameter | Type | Description |
|---|---|---|
text | char* | Text to synthesize. SSML is not supported. SSML tags in the input are read as plain text rather than parsed. |
stopStreamInputTts
Synchronous method. Notifies the server that all text has been sent, then blocks until every audio chunk is synthesized and TTS_EVENT_SYNTHESIS_COMPLETE is received.
The block timeout is controlled by complete_waiting_ms.
Method signature:
asyncStopStreamInputTts
Asynchronous method. Notifies the server that all text has been sent and returns immediately. Synthesis continues in the background.
Use TTS_EVENT_SYNTHESIS_COMPLETE to detect when synthesis finishes.
Method signature:
cancelStreamInputTts
Immediately drops the connection to the server and terminates the current synthesis task. After this method is called, no further audio data callbacks fire.
Method signature:
playStreamInputTts
Synchronous one-shot synthesis method. Sends the text and blocks while all audio data is received, then returns after synthesis finishes. You do not need to call stopStreamInputTts afterward.
This method enables SSML by default. To disable SSML, set the enable_ssml field in parameters to false.
Method signature:
ticket, parameters, and other shared parameters use the same definitions as startStreamInputTts.
| Parameter | Type | Description |
|---|---|---|
text | char* | Text to synthesize. Supports SSML. |
asyncPlayStreamInputTts
This method sends all text for synthesis asynchronously. It returns immediately without waiting for audio data. You do not need to call stopStreamInputTts afterward.
This method enables SSML by default. To disable SSML, set the enable_ssml field in parameters to false.
Method signature:
ticket, parameters, and other shared parameters use the same definitions as startStreamInputTts.
| Parameter | Type | Description |
|---|---|---|
text | char* | Text to synthesize. Supports SSML. |
StreamInputTtsDelegate
Callback protocol for Qwen-Audio-TTS/CosyVoice streaming speech synthesis. Implement this protocol to receive synthesis events, audio data, and logs.
onStreamInputTtsEventCallback: listen for events
Method signature:
| Parameter | Type | Description |
|---|---|---|
event | StreamInputTtsCallbackEvent | Callback event. |
taskid | char* | Speech synthesis task ID. |
sessionId | char* | Session ID. The client-supplied value is returned as-is. If none was supplied, the server generates one. |
ret_code | int | Error code. Valid only for the TTS_EVENT_TASK_FAILED event. See error codes. |
error_msg | char* | Error message. Valid only for the TTS_EVENT_TASK_FAILED event. |
timestamp | char* | Timestamp information for the synthesis result. |
all_response | char* | Full JSON response. Parse this string to extract the fields you need. |
onStreamInputTtsDataCallback: listen for audio data
The SDK fires this callback repeatedly during synthesis. Read the audio data from the callback.
Method signature:
| Parameter | Type | Description |
|---|---|---|
buffer | char* | Audio data for the current segment. Use this data to:
Note:
|
len | int | Length of the audio data, in bytes. |
onStreamInputTtsLogTrackCallback: listen for trace logs
This callback delivers detailed SDK-internal logs to help with troubleshooting and debugging.
Method signature:
| Parameter | Type | Description |
|---|---|---|
level | NuiSdkLogLevel | Log level. |
log | char* | Log content. |
StreamInputTtsCallbackEvent
Event-type enum for Qwen-Audio-TTS/CosyVoice streaming speech synthesis.
| Event | Description |
|---|---|
| TTS_EVENT_SYNTHESIS_STARTED | The server accepted the request and started processing. onStreamInputTtsDataCallback typically delivers the first audio segment shortly after this event. |
| TTS_EVENT_SENTENCE_SYNTHESIS | Progress information emitted during synthesis, including billing data. |
| TTS_EVENT_SYNTHESIS_COMPLETE | The server has finished sending all audio data. onStreamInputTtsDataCallback is not called again. This event is the definitive end-of-stream signal. |
| TTS_EVENT_TASK_FAILED | The task failed. Read task_id, error_code, and error_message from all_response in onStreamInputTtsEventCallback to diagnose the failure. |
NuiSdkLogLevel
SDK log-level enum that controls log output.
| Level | Description |
|---|---|
| 0: LOG_LEVEL_VERBOSE | Most detailed log level. Includes all debug information. |
| 1: LOG_LEVEL_DEBUG | Debug-level logs. |
| 2: LOG_LEVEL_INFO | General informational logs (default). |
| 3: LOG_LEVEL_WARNING | Warning-level logs. |
| 4: LOG_LEVEL_ERROR | Error-level logs. |
| 5: LOG_LEVEL_NONE | Disables log output. |
Sample code
-
Get your API key: Obtain an API key.
For third-party apps or end users that need temporary access, or when you want tight control over sensitive operations such as data access and deletion, use a temporary API key instead. A temporary API key is valid for a fixed 60 seconds and must be regenerated after it expires.
-
Download the SDK and run the sample code:
- Download the latest SDK bundle.
- Extract the ZIP archive and add
nuisdk.frameworkto your Xcode project. - In Build Phases > Link Binary With Libraries, add
nuisdk.framework. - In General > Frameworks, Libraries, and Embedded Content, set
nuisdk.frameworkto Embed & Sign. - Open the sample project in Xcode. The sample code is in
DashQwen-Audio-TTS/CosyVoiceStreamInputTTSViewController.m. Replace the placeholder API key with your own and run the app to try it out.
Call modes
| Call mode | Description |
|---|---|
| One-shot input | Steps:
Use cases:
|
| Streaming input | Steps:
Use cases:
|
Advanced features
SSML markup
Purpose: Embed XML tags in the input text to precisely control pronunciation, speech rate, pauses, and other prosody details.
Limitations: SSML is supported only with one-shot input (the playStreamInputTts and asyncPlayStreamInputTts methods). It is not supported with streaming input (the sendStreamInputTts method).
How to use: When you call playStreamInputTts or asyncPlayStreamInputTts, the SDK enables SSML automatically. Pass text that contains SSML tags directly in the text parameter.
For more details, see SSML.
Math expressions
Purpose: Make the model read common math formulas and expressions correctly.
How to use: Pass text that contains math expressions in LaTeX format directly in the text parameter. For more details, see Convert LaTeX formulas to speech (Chinese language only).