Qwen real-time speech and audiovisual translation with 2.3-second latency
Model details
Qwen real-time speech and audiovisual translation uses audio and image input to translate in real time and outputs text or speech in the target language for voice communication and video translation. The qwen3.8-livetranslate-flash-realtime and qwen3.5-livetranslate-flash-realtime models are vision-enhanced real-time translation models supporting 60 languages (29 with audio + text, 31 text-only). They process audio and image input from video streams or local files, use visual context to improve accuracy, and output translated text and audio in real time.
Key features:
- Multi-language support: Translates between 60 languages — 29 with audio and text output, 31 with text-only output — including Chinese, English, French, German, Russian, Japanese, Korean, Spanish, Portuguese, and Arabic.
- Visual enhancement: Analyzes visual cues, such as lip movements, gestures, and on-screen text, to improve translation accuracy, especially in noisy environments or for ambiguous words.
- 2.3-second latency: Delivers simultaneous interpretation with latency as low as 2.3 seconds.
- Real-time speaker diarization: Distinguishes different speakers and their speech when multiple people take turns speaking, so listeners can clearly understand who said what.
- Lossless simultaneous interpretation: Predicts semantic units to resolve cross-language word order differences, achieving quality comparable to offline translation.
- Natural voice: Matches the intonation and emotion of the source audio automatically.
- Hotword configuration: Configurable hotwords improve translation accuracy for specific terms.
- Voice cloning: Clones the speaker's voice for translated output. Supports server-side real-time cloning and pre-cloned voice profiles.
qwen3.5-livetranslate-flash-realtime also supports the AOQ and WebRTC protocols. For client-side integration that prioritizes stable latency, resilience on weak networks, and built-in full-duplex noise suppression and echo cancellation, AOQ is recommended. For a protocol comparison, see Realtime API overview.
Recommended models
| Model | Version | Context window | Max input | Max output |
|---|---|---|---|---|
| qwen3.8-livetranslate-flash-realtime | Stable | 53,248 | 49,152 | 4,096 |
| qwen3.5-livetranslate-flash-realtime (Alias for qwen3.5-livetranslate-flash-realtime-2026-05-19) | Stable | 53,248 | 49,152 | 4,096 |
| qwen3.5-livetranslate-flash-realtime-2026-05-19 | Snapshot | 53,248 | 49,152 | 4,096 |
Legacy models
The following model is still available but is no longer the recommended choice. For new use cases, use the newer model above for better translation quality and cost-efficiency.
| Model | Version | Context window | Max input | Max output |
|---|---|---|---|---|
| qwen3-livetranslate-flash-realtime (Alias for qwen3-livetranslate-flash-realtime-2025-09-22) | Stable | 53,248 | 49,152 | 4,096 |
| qwen3-livetranslate-flash-realtime-2025-09-22 | Snapshot | 53,248 | 49,152 | 4,096 |
Getting started
Prepare the environment
Requires Python 3.10 or later.
First, install pyaudio.
- macOS
- Debian/Ubuntu
- CentOS
- Windows
Create the client
Create a file named livetranslate_client.py with the following code:
Client code - livetranslate_client.py
Client code - livetranslate_client.py
Interact with the model
In the same directory, create a file named main.py with the following code:
main.py
main.py
main.py and speak into your microphone. The model outputs translated audio and text in real time. The system automatically detects speech and sends it to the server.
How to use
1. Configure the connection
The qwen3.5-livetranslate-flash-realtime model uses the WebSocket protocol. The connection requires the following parameters:
| Parameter | Description |
|---|---|
| endpoint | wss://maas.qwencloudapi.com/api-ws/v1/realtime |
| query parameter | The model query parameter must be set to the model name. Example: ?model=qwen3.5-livetranslate-flash-realtime |
| message header | Use a Bearer Token for authentication: Authorization: Bearer DASHSCOPE_API_KEY |
DASHSCOPE_API_KEY is your API key from QwenCloud.
Python sample code for WebSocket connection
Python sample code for WebSocket connection
2. Configure language, modality, and voice
- qwen3.8-livetranslate-flash-realtime
- qwen3.5-livetranslate-flash-realtime
Configure the session through session.update:
- Target language: Set
session.translation.language, such asenfor English. - Output modalities: Set
session.output_modalitiesto["text"]for text only or["text", "audio"]for text and audio. - Source transcription: Receive increments through
conversation.item.input_audio_transcription.deltaand complete transcripts throughconversation.item.input_audio_transcription.completed. - Audio and voice: Defaults are 16000 Hz PCM input, 24000 Hz PCM output, and voice
Tina. See Server events for the session structure.
3. Input audio and images
Send Base64-encoded audio and image data using the input_audio_buffer.append and input_image_buffer.append events. Audio input is required; image input is optional.
Images can be from a local file or captured in real time from a video stream.
For
qwen3.8-livetranslate-flash-realtime, the default turn detection configuration is audio.input.turn_detection.type = speaker_detection; send audio continuously and receive server-generated responses. The VAD and Manual configurations below apply to qwen3.5-livetranslate-flash-realtime.- VAD mode (default): The client continuously sends input_audio_buffer.append events. When the server detects speech start/end, it returns
input_audio_buffer.speech_startedandinput_audio_buffer.speech_stoppedevents respectively, automatically commits the audio buffer, and triggers translation. Translation responses are generated synchronously with the streaming audio and typically begin during audio input, without waiting for the speech to end. - Manual mode: Set
session.turn_detectiontonull. After the client finishes sending a complete utterance, it sends an input_audio_buffer.commit event to commit the audio buffer. After the server returns aninput_audio_buffer.committedevent to confirm, it automatically starts generating the translation response; the client does not need to send any other event to trigger the response. To clear uncommitted audio before committing, send an input_audio_buffer.clear event.
4. Receive the model response
- qwen3.8-livetranslate-flash-realtime
- qwen3.5-livetranslate-flash-realtime
Handle responses according to the output modalities:
- Text only: Concatenate the
deltavalues fromresponse.text.deltato obtain the translation. - Text and audio: Concatenate
deltafromresponse.audio_transcript.deltafor text, and Base64-decodedeltafromresponse.audio.deltato obtain audio chunks.
response.done indicates the end of a response. See Server events for event fields.5. End the session
After sending all audio, send a session.finish event, then wait for the server to return a session.finished event before closing the WebSocket connection.
If you close the WebSocket without sending
session.finish, the server's VAD cannot detect the end of the final speech segment. This causes translation results for that segment to be lost entirely, and the connection may hang indefinitely. Always send this event before disconnecting.Voice cloning
The model clones the speaker's voice from the input audio and uses the cloned voice for translated output, so the translation sounds like the speaker delivering it in another language. Use a pre-cloned voice profile, or let the server clone the voice in real time. This is useful in scenarios where preserving the speaker's voice matters, such as conference interpreting, live streaming, and video dubbing.
Set the following parameters in session.update to enable voice cloning:
session.enable_voice_clone: Set totrueto enable voice cloning.session.voice_clone_options.frequency: Controls when voice cloning occurs. Accepted values:never: Does not clone on the server. Uses a pre-cloned voice profile instead. Setsession.voiceto your custom cloned voice ID.once: Clones the voice from the input audio once at session start, then reuses it for all subsequent output. Best for single-speaker scenarios. Setsession.voicetodefault.always: Clones the voice before each response, dynamically adapting to speaker changes. Best for multi-speaker conversations. Setsession.voicetodefault.
session.voice: Specifies the output voice. The value depends on thefrequencysetting:- Set to
default: Use withfrequencyset toonceoralways. The server clones the speaker's voice from the input audio. A default voice is used until cloning completes. - Set to a custom cloned voice ID (for example,
qwen-translate-vc-xxx-yyy-zzz): Use withfrequencyset tonever. You must prepare the voice in advance using the Voice Cloning API withtargetModelset to the translation model you use.
- Set to
When
frequency is set to once or always, the voice parameter must be set to default. Any other value causes the server to return an error.Voice cloning configuration examples
Pre-cloned voice profile (consistent quality; recommended when a stable voice identity is required):
Interaction flow
- qwen3.8-livetranslate-flash-realtime
- qwen3.5-livetranslate-flash-realtime
The server detects turns and generates responses by default. The following table summarizes the main interaction stages. See Server events for event fields.
| Stage | Client action | Server events |
|---|---|---|
| Create and configure the session | Connect and send session.update | session.created, session.updated |
| Send audio | input_audio_buffer.append | Source transcription is streamed through conversation.item.input_audio_transcription.delta, followed by conversation.item.input_audio_transcription.completed. |
| Receive translation and audio | Keep receiving server events | Translation is returned through response.text.delta (text only) or response.audio_transcript.delta (text and audio). Audio is returned through response.audio.delta. response.done marks the end of a response. |
| End the session | session.finish | Close the connection after receiving session.finished. |
Improve translation with images
The qwen3.5-livetranslate-flash-realtime model uses image input to improve audio translation, helping disambiguate homonyms and recognize uncommon proper nouns. Send no more than 2 images per second.
Download the following sample images: medical mask.png and masquerade mask.png
Download the following code to the same directory as livetranslate_client.py and run it. Say "What is mask?" into your microphone. The model uses the provided image to disambiguate the word "mask." For example, using the medical mask.png file translates the phrase as "What is a medical mask?", while using the masquerade mask.png file translates it as "What is a masquerade mask?".
Billing
Qwen3.8-LiveTranslate-Flash-Realtime and Qwen3.5-LiveTranslate-Flash-Realtime
- Audio: 7 tokens per second of input audio; 12.5 tokens per second of output audio.
- Image: Every 32x32 pixels consumes 0.5 tokens.
- Audio: Each second of audio input or output consumes 12.5 tokens.
- Image: Every 28x28 pixels consumes 0.5 tokens.
- Text: When source language speech recognition is enabled, the service returns a transcript of the input audio in addition to the translation. This transcript is billed as output text tokens.
Supported languages
Use the following language codes to specify the source and target languages.
Some target languages only support text. The legacy model qwen3-livetranslate-flash-realtime supports only the following 18 languages: en, zh, ru, fr, de, pt, es, it, id, ko, ja, vi, th, ar, yue, hi, el, tr.
| Language code | Language | Output |
|---|---|---|
| zh | Chinese | Audio + text |
| en | English | Audio + text |
| ar | Arabic | Audio + text |
| de | German | Audio + text |
| fr | French | Audio + text |
| es | Spanish | Audio + text |
| pt | Portuguese | Audio + text |
| id | Indonesian | Audio + text |
| it | Italian | Audio + text |
| ko | Korean | Audio + text |
| ru | Russian | Audio + text |
| th | Thai | Audio + text |
| vi | Vietnamese | Audio + text |
| ja | Japanese | Audio + text |
| tr | Turkish | Audio + text |
| hi | Hindi | Audio + text |
| ms | Malay | Audio + text |
| nl | Dutch | Audio + text |
| ur | Urdu | Audio + text |
| nb | Norwegian Bokmål | Audio + text |
| sv | Swedish | Audio + text |
| da | Danish | Audio + text |
| he | Hebrew | Audio + text |
| fi | Finnish | Audio + text |
| pl | Polish | Audio + text |
| is | Icelandic | Audio + text |
| cs | Czech | Audio + text |
| fil | Filipino | Audio + text |
| fa | Persian | Audio + text |
| yue | Cantonese | Text |
| el | Greek | Text |
| af | Afrikaans | Text |
| ast | Asturian | Text |
| be | Belarusian | Text |
| bg | Bulgarian | Text |
| bn | Bengali | Text |
| bs | Bosnian | Text |
| ca | Catalan | Text |
| ceb | Cebuano | Text |
| et | Estonian | Text |
| gl | Galician | Text |
| gu | Gujarati | Text |
| hr | Croatian | Text |
| hu | Hungarian | Text |
| jv | Javanese | Text |
| kk | Kazakh | Text |
| kn | Kannada | Text |
| ky | Kyrgyz | Text |
| lv | Latvian | Text |
| mk | Macedonian | Text |
| ml | Malayalam | Text |
| mr | Marathi | Text |
| pa | Punjabi | Text |
| ro | Romanian | Text |
| sk | Slovak | Text |
| sl | Slovenian | Text |
| sw | Swahili | Text |
| tg | Tajik | Text |
| az | Azerbaijani | Text |
| uk | Ukrainian | Text |
Supported voices
For the voices supported by Qwen3.5-LiveTranslate and Qwen3-LiveTranslate, and their voice parameter values, see Omni-modal voice list.
API reference
- Client events
- Server events
- Python SDK
- Java SDK
- AOQ client SDK
- Realtime API overview (WebRTC protocol description)