Choose a model for live captions, file transcription, and more.
Migrate from closed-source models
Replacing Whisper, Deepgram, or Google speech recognition? Use these QwenCloud equivalents.
| Use case | Closed-source examples | QwenCloud recommendation |
|---|---|---|
| Real-time recognition | Deepgram Nova-3, Google Chirp 3 | qwen-audio-3.0-asr-flash-streaming |
| Offline / file transcription | OpenAI gpt-4o-transcribe, Whisper | qwen-audio-3.0-asr-flash-filetrans, qwen-audio-3.0-asr-flash |
Decision dimensions
Use these 4 dimensions to narrow your choice. Each recommends the best-fit model.
Real-time or offline?
Real-time means outputting recognition results while the user is still speaking. Offline means transcribing after the recording ends.
-
Real-time (streaming recognition) — Uses a WebSocket connection to stream audio in and text out. Ideal for live captions, voice assistants, and meeting transcription. Recommended model:
qwen-audio-3.0-asr-flash-streaming(hot words, prompt context, multilingual with dialects). -
Offline (file transcription) — Uses an HTTP API to submit audio files and retrieve transcription results. Suited for call center recordings, podcasts, and interviews. Recommended model:
qwen-audio-3.0-asr-flash-filetrans(hot words, prompt context, speaker diarization).
Handling domain terminology
Two approaches, ranked by flexibility:
- Prompt context injection — Describe your domain in the system prompt. No setup required -- the model adapts per request. Recommended: the Qwen-Audio-3.0-ASR-Flash-Streaming, Qwen-Audio-3.0-ASR-Flash-Filetrans, and Qwen-Audio-3.0-ASR-Flash series.
- Hot words — Supply a weighted vocabulary list. Recommended: the Qwen-Audio-3.0-ASR-Flash-Streaming, Qwen-Audio-3.0-ASR-Flash-Filetrans, and Qwen-Audio-3.0-ASR-Flash series.
Speaker diarization
The Qwen-Audio-3.0-ASR-Flash-Filetrans series (qwen-audio-3.0-asr-flash-filetrans) and Fun-ASR offline models (fun-asr, fun-asr-mtl) support speaker diarization. To identify who said what, use qwen-audio-3.0-asr-flash-filetrans.
Emotion recognition
Qwen-ASR models detect emotions alongside transcription. Recommended models: qwen3-asr-flash-realtime (real-time) or qwen3-asr-flash-filetrans (offline).
Recommended models
The top picks for each scenario. Visit the model gallery for details.
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
qwen-audio-3.0-asr-flash-streaming | Real-time | WebSocket | Hot words, prompt context | ✗ | ✗ | Multilingual with dialects | Unlimited |
qwen-audio-3.0-asr-flash-filetrans | Offline | HTTP | Hot words, prompt context | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
All models
Qwen-Audio-3.0-ASR-Flash-Streaming
Qwen-Audio-3.0-ASR-Flash-Streaming
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
qwen-audio-3.0-asr-flash-streaming | Real-time | WebSocket | Hot words, prompt context | ✗ | ✗ | Multilingual with dialects | Unlimited |
Qwen-Audio-3.0-ASR-Flash-Filetrans
Qwen-Audio-3.0-ASR-Flash-Filetrans
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
qwen-audio-3.0-asr-flash-filetrans | Offline | HTTP | Hot words, prompt context | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
Qwen-Audio-3.0-ASR-Flash
Qwen-Audio-3.0-ASR-Flash
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
qwen-audio-3.0-asr-flash | Offline | HTTP | Hot words, prompt context | ✗ | ✗ | Multilingual with dialects | 5 min / 2GB |
Fun-ASR
Fun-ASR
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
fun-asr-realtime | Real-time | WebSocket | Hot words | ✗ | ✗ | Multilingual with dialects | Unlimited |
fun-asr-realtime-2026-02-28 | Real-time | WebSocket | Hot words | ✗ | ✗ | Chinese, English, Japanese | Unlimited |
fun-asr-realtime-2025-11-07 | Real-time | WebSocket | Hot words | ✗ | ✗ | Multilingual with dialects | Unlimited |
fun-asr-realtime-2025-09-15 | Real-time | WebSocket | Hot words | ✗ | ✗ | Chinese, English | Unlimited |
fun-asr-flash-8k-realtime | Real-time | WebSocket | Hot words | ✗ | ✗ | Chinese | Unlimited |
fun-asr-flash-8k-realtime-2026-01-28 | Real-time | WebSocket | Hot words | ✗ | ✗ | Chinese | Unlimited |
fun-asr | Offline | HTTP | Hot words | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
fun-asr-2025-11-07 | Offline | HTTP | Hot words | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
fun-asr-2025-08-25 | Offline | HTTP | Hot words | ✗ | ✓ | Chinese, English | 12h / 2GB |
fun-asr-mtl | Offline | HTTP | Hot words | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
fun-asr-mtl-2025-08-25 | Offline | HTTP | Hot words | ✗ | ✓ | Multilingual with dialects | 12h / 2GB |
fun-asr-flash-2026-06-15 | Offline | HTTP | Prompt context | ✗ | ✗ | Multilingual with dialects | 5 min / 2GB |
- Fun-ASR-Realtime main versions (
fun-asr-realtime,fun-asr-realtime-2026-02-28,fun-asr-realtime-2025-11-07): Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin; plus regional accents including Central Plains, Southwestern, Ji-Lu, Jianghuai, Lan-Yin, Jiao-Liao, Northeastern, Beijing, and Hong Kong/Taiwan -- covering Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, Hebei, Tianjin, Shandong, Anhui, Nanjing, Jiangsu, Hangzhou, Gansu, and Ningxia accents), English, Japanese fun-asr-realtime-2025-09-15: Chinese (Mandarin), English- Fun-ASR-Flash-Realtime(8K) (
fun-asr-flash-8k-realtime,fun-asr-flash-8k-realtime-2026-01-28): Chinese - Fun-ASR main versions (
fun-asr,fun-asr-2025-11-07): Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin; plus regional accents including Central Plains, Southwestern, Ji-Lu, Jianghuai, Lan-Yin, Jiao-Liao, Northeastern, Beijing, and Hong Kong/Taiwan -- covering Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, Hebei, Tianjin, Shandong, Anhui, Nanjing, Jiangsu, Hangzhou, Gansu, and Ningxia accents), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak - Fun-ASR-Flash (
fun-asr-flash-2026-06-15): Chinese (Mandarin, Cantonese, Wu, Hokkien, Hakka, Gan, Xiang, Jin; plus regional accents including Central Plains, Southwestern, Ji-Lu, Jianghuai, Lan-Yin, Jiao-Liao, Northeastern, Beijing, and Hong Kong/Taiwan -- covering Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, Hebei, Tianjin, Shandong, Anhui, Nanjing, Jiangsu, Hangzhou, Gansu, and Ningxia accents), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak fun-asr-2025-08-25: Chinese (Mandarin), English- Fun-ASR(MTL) (
fun-asr-mtl,fun-asr-mtl-2025-08-25): Chinese (Mandarin, Cantonese), English, Japanese, Korean, Vietnamese, Thai, Indonesian, Malay, Filipino, Hindi, Arabic, French, German, Spanish, Portuguese, Russian, Italian, Dutch, Swedish, Danish, Finnish, Norwegian, Greek, Polish, Czech, Hungarian, Romanian, Bulgarian, Croatian, Slovak
Qwen-ASR
Qwen-ASR
| Model | Mode | API | Accuracy boost | Emotion | Speaker diarization | Languages | Max duration/size |
|---|---|---|---|---|---|---|---|
qwen3-asr-flash-realtime | Real-time | WebSocket | ✗ | ✓ | ✗ | Multilingual with dialects | Unlimited |
qwen3-asr-flash-realtime-2026-02-10 | Real-time | WebSocket | ✗ | ✓ | ✗ | Multilingual with dialects | Unlimited |
qwen3-asr-flash-realtime-2025-10-27 | Real-time | WebSocket | ✗ | ✓ | ✗ | Multilingual with dialects | Unlimited |
qwen3-asr-flash-filetrans | Offline | HTTP | ✗ | ✓ | ✗ | Multilingual with dialects | 12h / 2GB |
qwen3-asr-flash-filetrans-2025-11-17 | Offline | HTTP | ✗ | ✓ | ✗ | Multilingual with dialects | 12h / 2GB |
qwen3-asr-flash | Offline | HTTP (OpenAI compatible) | ✗ | ✓ | ✗ | Multilingual with dialects | 5 min / 10MB |
qwen3-asr-flash-2026-02-10 | Offline | HTTP (OpenAI compatible) | ✗ | ✓ | ✗ | Multilingual with dialects | 5 min / 10MB |
qwen3-asr-flash-2025-09-08 | Offline | HTTP (OpenAI compatible) | ✗ | ✓ | ✗ | Multilingual with dialects | 5 min / 10MB |
qwen3-asr-flash-realtime, qwen3-asr-flash-filetrans, qwen3-asr-flash and their snapshot versions) support the same language list: Chinese (Mandarin, Sichuanese, Hokkien, Wu, Cantonese), English, Japanese, German, Korean, Russian, French, Portuguese, Arabic, Italian, Spanish, Hindi, Indonesian, Thai, Turkish, Ukrainian, Vietnamese, Czech, Danish, Filipino, Finnish, Icelandic, Malay, Norwegian, Polish, Swedish.Paraformer
Paraformer
Paraformer is an earlier-generation ASR model family. If your workload allows, migrate to the recommended Fun-ASR or Qwen-ASR models above.
Supported languages (by version):
| Model | API | Description |
|---|---|---|
paraformer-realtime-v2 | WebSocket | Real-time recognition; Chinese, English, Japanese, Korean, German, French, Russian |
paraformer-realtime-v1 | WebSocket | Real-time recognition; Chinese, English, Japanese, Korean, German, French, Russian |
paraformer-realtime-8k-v2 | WebSocket | Real-time recognition; 8 kHz telephony; Chinese |
paraformer-realtime-8k-v1 | WebSocket | Real-time recognition; 8 kHz telephony; Chinese |
paraformer-v2 | HTTP | File transcription with speaker diarization; Chinese, English, Japanese, Korean, German, French, Russian |
paraformer-8k-v2 | HTTP | File transcription; 8 kHz telephony; Chinese |
paraformer-realtime-v2,paraformer-v2: Chinese (Mandarin, Cantonese, Wu, Hokkien, Northeastern, Gansu, Guizhou, Henan, Hubei, Hunan, Ningxia, Shanxi, Shaanxi, Shandong, Sichuan, Tianjin, Jiangxi, Yunnan, Shanghai), English, Japanese, Korean, German, French, Russianparaformer-realtime-v1,paraformer-realtime-8k-v2,paraformer-realtime-8k-v1,paraformer-8k-v2: Chinese (Mandarin)
Audio specifications
The tables below summarize audio specs (input method, format, sample rate, size/duration) for real-time and offline modes. For supported languages, see the model family sections under "All models" above.
Real-time
| Model | Input method | Audio format | Sample rate | Size/duration |
|---|---|---|---|---|
Qwen-Audio-3.0-ASR-Flash-Streaming (qwen-audio-3.0-asr-flash-streaming) | Binary stream | pcm, wav, mp3, opus, speex, aac, amr | Any | Unlimited |
Fun-ASR-Realtime (fun-asr-realtime family) | Binary stream | pcm, wav, mp3, opus, speex, aac, amr | Any | Unlimited |
Fun-ASR-Flash-8K-Realtime (fun-asr-flash-8k-realtime family) | Binary stream | Same as Fun-ASR-Realtime | 8 kHz | Unlimited |
Qwen-ASR-Realtime (qwen3-asr-flash-realtime family) | Binary stream | pcm, opus | 8 kHz, 16 kHz | Unlimited |
Paraformer-Realtime (paraformer-realtime-v2/v1, paraformer-realtime-8k-v2/v1) | Binary stream | Same as Fun-ASR-Realtime | paraformer-realtime-v2 any; paraformer-realtime-v1 16 kHz; paraformer-realtime-8k-* 8 kHz | Unlimited |
Offline
| Model | Input method | Audio format | Sample rate | File size/duration |
|---|---|---|---|---|
Qwen-Audio-3.0-ASR-Flash-Filetrans (qwen-audio-3.0-asr-flash-filetrans) | Publicly accessible file URL, 1 per request | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv | Any | ≤2 GB; ≤12 hours (≤2 hours recommended with speaker diarization) |
Qwen-Audio-3.0-ASR-Flash (qwen-audio-3.0-asr-flash) | URL / Base64, 1 per request | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv | Any | ≤2 GB; ≤5 min |
Fun-ASR (fun-asr, fun-asr-mtl family) | Publicly accessible file URL, 1 per request | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv | Any | ≤2 GB; ≤12 hours (≤2 hours recommended with speaker diarization) |
Fun-ASR-Flash (fun-asr-flash-2026-06-15) | URL / Base64, 1 per request | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv | Any | ≤2 GB; ≤5 min |
Paraformer (paraformer-v2/v1, paraformer-mtl-v1, paraformer-8k-v2/v1) | Same as Fun-ASR | Same as Fun-ASR | paraformer-v2 any; paraformer-8k-* 8 kHz only | Same as Fun-ASR |
Qwen3-ASR-Flash-Filetrans (qwen3-asr-flash-filetrans family) | Publicly accessible file URL, 1 per request | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv | pcm must be 16 kHz; other formats any (server resamples to 16 kHz) | ≤2 GB; ≤12 hours |
Qwen3-ASR-Flash (qwen3-asr-flash family) | URL / Base64 / local file absolute path, 1 per request | aac, amr, avi, aiff, flac, flv, mkv, mp3, mpeg, ogg, opus, wav, webm, wma, wmv | pcm must be 16 kHz; other formats any (server resamples to 16 kHz) | ≤10 MB; ≤5 min |
Need speech translated to a different language? See Speech-to-Speech models for real-time and file-based translation with LiveTranslate and Qwen-Omni.