Convert files to text
Non-real-time speech recognition models convert recorded audio into text. They support multilingual recognition, singing recognition, noise rejection, and speaker diarization, which makes them suitable for meeting transcription, call analysis, subtitle generation, and similar scenarios.
QwenCloud also offers Qwen-ASR for recognition with enhanced semantic understanding and Qwen-Omni for prompt-based transcription with contextual understanding.
Supported models: Qwen-Audio-3.x-ASR-Flash (
To produce the corrected result, include any of the following in the context:
In addition to context enhancement, audio file transcription supports the following capabilities. To create and use hotwords, see Improve recognition accuracy.
Audio file transcription supports asynchronous transcription of long audio, which suits meeting minutes, interview notes, and call playback.
Limits:
Speaker diarization identifies different speakers in the audio and labels each sentence in the transcription result, which suits multi-person meetings and interview recordings.
Qwen-Audio-3.1-ASR-Flash enables speaker diarization with
SDKs expose these fields with different naming conventions (dictionary keys, object properties, or methods). For the full field mapping, see API reference.
Sensitive word filtering replaces or removes sensitive words in the recognition result, which suits customer service quality checks, content compliance, and subtitle review.
Supported models: Qwen-Audio-3.x-ASR-Flash-Filetrans, Fun-ASR, and Paraformer.
Default behavior: If you do not pass the
Qwen3-ASR-Flash-Filetrans and Qwen3-ASR-Flash always perform emotion recognition, with no extra configuration. The result includes the speaker's emotion label, which is one of seven values:
Audio file transcription can return timestamps in the result, which suits subtitle generation, keyword highlighting, and audio or video editing. The default behavior and controls vary by model:
Timestamps are in milliseconds and are returned at two levels:
When you move audio file transcription to production, the following practices improve recognition quality and system stability.
For model availability, supported languages, and feature comparison, see Speech-to-text models.
Core features
- Multilingual recognition: Recognizes Chinese (including multiple dialects), English, Japanese, Korean, German, French, Russian, and 30+ other languages.
- Format compatibility: Accepts any sample rate and supports major audio and video formats, including AAC, WAV, and MP3.
- Long audio file processing: Handles asynchronous transcription for a single audio file up to 12 hours long and 2 GB in size. If speaker diarization is enabled, audio longer than 2 hours is not recommended.
- Singing voice recognition: Transcribes entire songs, even with background music (BGM). Only the fun-asr and fun-asr-2025-11-07 models support this feature.
- Recognition features: Configurable features include speaker diarization, sensitive word filtering, sentence-level and word-level timestamps, and hotword enhancement.
Supported models
- Qwen-Audio-3.1-ASR-Flash-Filetrans: qwen-audio-3.1-asr-flash-filetrans
- Qwen-Audio-3.0-ASR-Flash-Filetrans: qwen-audio-3.0-asr-flash-filetrans
- Qwen-Audio-3.1-ASR-Flash: qwen-audio-3.1-asr-flash
- Qwen-Audio-3.0-ASR-Flash: qwen-audio-3.0-asr-flash
- Fun-ASR: fun-asr (stable, currently equivalent to fun-asr-2025-11-07), fun-asr-2025-11-07 (snapshot), fun-asr-2025-08-25 (snapshot), fun-asr-mtl (stable, currently equivalent to fun-asr-mtl-2025-08-25, fun-asr is recommended), fun-asr-mtl-2025-08-25 (snapshot)
- Fun-ASR-Flash: fun-asr-flash-2026-06-15 (snapshot)
- Qwen3-ASR-Flash-Filetrans: qwen3-asr-flash-filetrans (stable, currently equivalent to qwen3-asr-flash-filetrans-2025-11-17), qwen3-asr-flash-filetrans-2025-11-17 (snapshot)
- Qwen3-ASR-Flash: qwen3-asr-flash (stable, currently equivalent to qwen3-asr-flash-2025-09-08), qwen3-asr-flash-2026-02-10 (latest snapshot), qwen3-asr-flash-2025-09-08 (snapshot)
Prerequisites
- Get an API key and set it as an environment variable.
- To call the API through the DashScope SDK, install the latest SDK.
Getting started
- Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR
- Qwen-ASR
- Qwen-Omni
Model availability
| Model | Version | Unit price | Free quota (Note) |
|---|---|---|---|
| fun-asr Currently, fun-asr-2025-11-07 | Stable | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
| fun-asr-2025-11-07 Improved far-field VAD over fun-asr-2025-08-25 for higher accuracy | Snapshot | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
| fun-asr-2025-08-25 | Snapshot | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
| fun-asr-mtl Currently, fun-asr-mtl-2025-08-25 | Stable | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
| fun-asr-mtl-2025-08-25 | Snapshot | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
| fun-asr-flash-2026-06-15 Supports synchronous calls (up to 5 minutes) and context enhancement | Snapshot | $0.000035/second | 36,000 seconds (10 hours) Valid for 90 days |
-
Supported languages:
- fun-asr, fun-asr-2025-11-07, fun-asr-mtl, and fun-asr-mtl-2025-08-25: 30 languages
- fun-asr-2025-08-25: Mandarin and English.
- Sample rates supported: Any
- Audio formats supported: aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv
Make your first call
Get an API key and set it as an environment variable. To use the SDK, install it.Because audio and video files are often large, file transfer and speech recognition can take a long time. The file recognition API uses asynchronous invocation to submit tasks. After the file recognition is complete, you must use the query API to retrieve the speech recognition results.Async submit and sync wait
Submit a task and block until done.The complete recognition result is printed to the console in JSON format. The result includes the transcribed text and the start and end times of the text in the audio or video file, specified in milliseconds.
The complete recognition result is printed to the console in JSON format. The result includes the transcribed text and the start and end times of the text in the audio or video file, specified in milliseconds.
First resultSecond result
Async submit and async query
Submit a task and poll for results instead of blocking.RESTful API
Use any HTTP library to submit tasks and poll for results. This Python sample demonstrates the workflow:Synchronous calls (Qwen-Audio-3.x-ASR-Flash/Fun-ASR-Flash)
The Qwen-Audio-3.x-ASR-Flash and Fun-ASR-Flash model series support synchronous calls for audio files shorter than 5 minutes. Results can be returned in streaming or non-streaming mode.Response structure note: The Qwen-Audio-3.x-ASR-Flash and Fun-ASR-Flash model series return a response structure through the DashScope synchronous API (the multimodal-generation endpoint) that differs from the standard DashScope multimodal response format. The recognized text is available at
output.output.sentence.text and output.text in the response, not at output.choices[].message.content. When parsing responses from these models, use the output.output path to access the recognition results.Example response excerpt:Context enhancement
Supported models: Qwen-Audio-3.x-ASR-Flash (qwen-audio-3.1-asr-flash) and Fun-ASR-Flash (fun-asr-flash-2026-06-15) support context enhancement.
Use case: Designed for scenarios that combine ASR with a large language model. Passing previous conversation context (LLM replies and earlier recognition results) into the ASR model significantly improves transcription accuracy for proper nouns such as names, locations, and product terms — more flexible than traditional hotwords.
Usage: Pass the conversation history through input.messages. Use the assistant role for the LLM's previous replies and the user role with input_text type for earlier recognition results. Context pairs must appear before the current audio message.
Supported text types include (but are not limited to):
- Hotword lists in various delimiter formats (for example: hotword1, hotword2, hotword3, hotword4)
- Free-form paragraphs or passages of any length
- Mixed content: any combination of word lists and paragraphs
- Irrelevant or meaningless text, including gibberish. The model tolerates irrelevant content well, and recognition quality rarely degrades because of it.
| Without context enhancement | With context enhancement |
|---|---|
| Without context enhancement, some investment bank names are misrecognized. For example, "Bird Rock" should be "Bulge Bracket". Result: "What insider jargon do you know in investment banking? First, the nine top foreign investment banks, Bird Rock, BB..." | With context enhancement, the investment bank names are recognized correctly. Result: "What insider jargon do you know in investment banking? First, the nine top foreign investment banks, Bulge Bracket, BB..." |
- A word list:
- List 1:
- List 2:
- List 3:
- Natural language:
- Natural language with distracting content: some text is unrelated to the audio, such as the names in the example below.
Advanced features
In addition to context enhancement, audio file transcription supports the following capabilities. To create and use hotwords, see Improve recognition accuracy.
Long audio file processing
Audio file transcription supports asynchronous transcription of long audio, which suits meeting minutes, interview notes, and call playback.
Limits:
- Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR/Qwen3-ASR-Flash-Filetrans: Each audio file must be no larger than 2 GB and no longer than 12 hours.
- Qwen-Audio-3.1-ASR-Flash: Each audio file must be no larger than 2 GB and no longer than 5 minutes.
- Qwen-Audio-3.x-ASR-Flash/Fun-ASR-Flash/Qwen3-ASR-Flash: Each audio file must be no larger than 10 MB and no longer than 5 minutes. For longer audio, use Qwen-Audio-3.1-ASR-Flash-Filetrans, Fun-ASR, or Qwen3-ASR-Flash-Filetrans.
- When speaker diarization is enabled: Audio longer than 2 hours is not recommended, because recognition may fail or time out. See Speaker diarization.
- Submit the transcription task and get a
task_id. - Poll for the task status, or use the SDK's wait method to block until the task finishes.
- When the task completes, download the result JSON from the returned URL.
Speaker diarization
Speaker diarization identifies different speakers in the audio and labels each sentence in the transcription result, which suits multi-person meetings and interview recordings.
Qwen-Audio-3.1-ASR-Flash enables speaker diarization with speaker_diarization_enabled. For its parameters and response structure, see RESTful API.
Supported models: Qwen-Audio-3.x-ASR-Flash-Filetrans, Fun-ASR, and Paraformer.
How to enable: Set diarization_enabled to true in the request parameters. Each sentence in the result then includes a speaker_id field that identifies the speaker.
Response structure (excerpt):
When speaker diarization is enabled, audio longer than 2 hours is not recommended, because recognition may fail or time out. For the duration limits that apply when diarization is off, see Long audio file processing. Speaker diarization supports mono audio only.
Sensitive word filtering
Sensitive word filtering replaces or removes sensitive words in the recognition result, which suits customer service quality checks, content compliance, and subtitle review.
Supported models: Qwen-Audio-3.x-ASR-Flash-Filetrans, Fun-ASR, and Paraformer.
Default behavior: If you do not pass the special_word_filter parameter, the built-in sensitive word list applies and matched words are replaced with an equal number of asterisks (*).
Custom configuration: special_word_filter is a JSON object with three sub-fields:
filter_with_signed.word_list: A string array of sensitive words to replace with an equal number of asterisks. For example, with["test"], "Run a test for me" becomes "Run a **** for me".filter_with_empty.word_list: A string array of sensitive words to remove entirely from the result. For example, with["start"], "Is the game about to start" becomes "Is the game about to".system_reserved_filter: A boolean that defaults totrue. Specifies whether to also apply the built-in sensitive word list, which is combined with your custom lists.
The default differs from real-time speech recognition: in real-time recognition, no filtering is applied when
special_word_filter is omitted, because system_reserved_filter defaults to false.Emotion recognition
Qwen3-ASR-Flash-Filetrans and Qwen3-ASR-Flash always perform emotion recognition, with no extra configuration. The result includes the speaker's emotion label, which is one of seven values: surprised, neutral, happy, sad, disgusted, angry, and fearful.
Field paths (they vary by interface):
- OpenAI-compatible interface (Qwen3-ASR-Flash):
choices[].delta.annotations[].emotionfor streaming output, orchoices[].message.annotations[].emotionfor non-streaming output. - DashScope synchronous interface (Qwen3-ASR-Flash):
output.choices[].message.annotations[].emotion. - DashScope asynchronous task interface (Qwen3-ASR-Flash-Filetrans):
transcripts[].sentences[].emotion, alongside the timestamp fields in each sentence object.
Qwen-Audio-3.x-ASR-Flash-Filetrans, Qwen-Audio-3.x-ASR-Flash, Fun-ASR-Flash, Fun-ASR, and Paraformer do not support emotion recognition. To use emotion recognition in real-time scenarios, see Real-time speech recognition.
Get timestamps
Audio file transcription can return timestamps in the result, which suits subtitle generation, keyword highlighting, and audio or video editing. The default behavior and controls vary by model:
- Qwen-Audio-3.x-ASR-Flash-Filetrans/Qwen-Audio-3.x-ASR-Flash/Fun-ASR/Fun-ASR-Flash: Timestamps are always on and cannot be turned off.
- Qwen3-ASR-Flash-Filetrans: Timestamps are supported only through the DashScope asynchronous interface, where they are always on. Use the
enable_wordsrequest parameter to control the granularity:false(default) returns sentence-level timestamps, andtruereturns word-level timestamps. Word-level timestamps are supported only for Chinese, English, Japanese, Korean, German, French, Spanish, Italian, Portuguese, and Russian. Accuracy is not guaranteed for other languages.
When you call Qwen3-ASR-Flash through the OpenAI-compatible interface, the output is a
chat.completion object and contains no timestamp fields. If you need timestamps, use Qwen3-ASR-Flash-Filetrans through the asynchronous task interface.- Sentence level:
sentences[].begin_timeandsentences[].end_timemark the start and end of each sentence in the audio. - Word level: The
sentences[].words[]array, where each element containsbegin_time,end_time, andtext(the word).
Timestamps inside the audio are millisecond integers such as
100. Do not confuse them with the task-level end_time, which is the task completion time as a date string such as "2024-09-12 15:11:40.903".Going live
When you move audio file transcription to production, the following practices improve recognition quality and system stability.
- File hosting: Upload audio files to an object storage service such as OSS and call the API by URL. Avoid local file uploads, which are capped at 100 QPS and cannot be scaled up.
- Asynchronous polling: Long audio transcription is asynchronous, and the task query API is limited to 20 QPS by default. Set a reasonable polling interval, such as 2 to 5 seconds, to avoid triggering rate limits. For polling strategies and batch task management, see Asynchronous task management.
- Error handling: Implement a robust retry mechanism. For network timeouts or temporary server-side errors (5xx), retry with an exponential backoff strategy.
- Noise reduction: For noisy audio, preprocess it with a tool such as FFmpeg before you submit it for recognition.
- Model selection: Choose the model based on audio duration. For audio within 5 minutes, use Qwen3-ASR-Flash. For audio longer than 5 minutes, use Qwen-Audio-3.1-ASR-Flash-Filetrans, Fun-ASR, or Qwen3-ASR-Flash-Filetrans.
Compare models
| Feature | Fun-ASR |
|---|---|
| Supported languages | Varies by model: fun-asr, fun-asr-2025-11-07, fun-asr-mtl, fun-asr-mtl-2025-08-25: Chinese (Mandarin, Cantonese, Wu, Minnan, Hakka, Gan, Xiang, Jin; also supports accents from Central Plains, Southwest, Ji-Lu, Jianghuai, Lan-Yin, Jiao-Liao, Northeast, Beijing, Hong Kong, and Taiwan, including official dialects from regions such as Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, Hebei, Tianjin, Shandong, Anhui, Nanjing, Jiangsu, Hangzhou, Gansu, and Ningxia), English, Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, Arabic, Bulgarian, Croatian, Czech, Danish, Dutch, Estonian, Finnish, Greek, Hindi, Hungarian, Irish, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish. fun-asr-2025-08-25: Chinese (Mandarin), English |
| Supported audio formats | aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv |
| Sample rate | Any |
| Sound channels | Any |
| Input format | Publicly accessible URLs of files to be recognized. Up to 100 audio files are supported. |
| Audio size/duration | Each audio file must be no larger than 2 GB and no longer than 12 hours. |
| Emotion recognition | Not supported |
| Timestamp | Supported (always on) |
| Punctuation prediction | Supported (always on) |
| Hotwords | Supported. The hotword feature is supported only in the primary workspace and is not available in sub-workspaces. |
| ITN | Supported (always on) |
| Singing voice recognition | Supported (fun-asr and fun-asr-2025-11-07 only) |
| Noise rejection | Supported (always on) |
| Sensitive word filtering | Supported (filters content from the QwenCloud sensitive word list by default) |
| Speaker diarization | Supported (off by default, can be enabled) |
| Filler word filtering | Not supported |
| VAD | Supported (always on) |
| Rate limiting (RPS) | Job submission API: 10, Task query API: 20 |
| Connection types | DashScope: Java/Python SDK, RESTful API |
| Pricing | International: $0.000035/second |
API reference
- Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR
- Qwen-ASR
- Qwen-Omni
FAQ
- Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR
- Qwen-ASR
- Qwen-Omni
How can I improve recognition accuracy?
Several factors affect accuracy. Review each and apply the corresponding optimization.Key factors:- Sound quality: Recording device quality, sample rate, and ambient noise directly affect clarity. High-quality audio input is essential.
- Speaker characteristics: Variations in pitch, speech rate, accent, and dialect increase recognition difficulty, especially for rare dialects or heavy accents.
- Language and vocabulary: Mixed languages, technical terms, or slang increase recognition difficulty. Configure hotwords to improve accuracy for domain-specific terms.
- Contextual understanding: Insufficient context can cause semantic ambiguity, especially in situations where surrounding context is needed for correct recognition.
- Optimize audio quality: Use high-performance microphones at the recommended sample rate. Minimize ambient noise and echo.
- Adapt to the speaker: For audio with strong accents or dialects, select a model that supports those specific dialects.
- Configure hotwords: Set hotwords for technical terms, proper nouns, and other specific words. For more information, see Customize hotwords.
- Preserve context: Avoid splitting audio into excessively short clips.