Skip to main content
Speech-to-text

Audio file transcription

Convert files to text

Non-real-time speech recognition models convert recorded audio into text. They support multilingual recognition, singing recognition, noise rejection, and speaker diarization, which makes them suitable for meeting transcription, call analysis, subtitle generation, and similar scenarios. QwenCloud also offers Qwen-ASR for recognition with enhanced semantic understanding and Qwen-Omni for prompt-based transcription with contextual understanding.
For model availability, supported languages, and feature comparison, see Speech-to-text models.

Core features

  • Multilingual recognition: Recognizes Chinese (including multiple dialects), English, Japanese, Korean, German, French, Russian, and 30+ other languages.
  • Format compatibility: Accepts any sample rate and supports major audio and video formats, including AAC, WAV, and MP3.
  • Long audio file processing: Handles asynchronous transcription for a single audio file up to 12 hours long and 2 GB in size. If speaker diarization is enabled, audio longer than 2 hours is not recommended.
  • Singing voice recognition: Transcribes entire songs, even with background music (BGM). Only the fun-asr and fun-asr-2025-11-07 models support this feature.
  • Recognition features: Configurable features include speaker diarization, sensitive word filtering, sentence-level and word-level timestamps, and hotword enhancement.

Supported models

  • Qwen-Audio-3.1-ASR-Flash-Filetrans: qwen-audio-3.1-asr-flash-filetrans
  • Qwen-Audio-3.0-ASR-Flash-Filetrans: qwen-audio-3.0-asr-flash-filetrans
  • Qwen-Audio-3.1-ASR-Flash: qwen-audio-3.1-asr-flash
  • Qwen-Audio-3.0-ASR-Flash: qwen-audio-3.0-asr-flash
  • Fun-ASR: fun-asr (stable, currently equivalent to fun-asr-2025-11-07), fun-asr-2025-11-07 (snapshot), fun-asr-2025-08-25 (snapshot), fun-asr-mtl (stable, currently equivalent to fun-asr-mtl-2025-08-25, fun-asr is recommended), fun-asr-mtl-2025-08-25 (snapshot)
  • Fun-ASR-Flash: fun-asr-flash-2026-06-15 (snapshot)
  • Qwen3-ASR-Flash-Filetrans: qwen3-asr-flash-filetrans (stable, currently equivalent to qwen3-asr-flash-filetrans-2025-11-17), qwen3-asr-flash-filetrans-2025-11-17 (snapshot)
  • Qwen3-ASR-Flash: qwen3-asr-flash (stable, currently equivalent to qwen3-asr-flash-2025-09-08), qwen3-asr-flash-2026-02-10 (latest snapshot), qwen3-asr-flash-2025-09-08 (snapshot)

Prerequisites

Getting started

  • Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR
  • Qwen-ASR
  • Qwen-Omni

Model availability

ModelVersionUnit priceFree quota (Note)
fun-asr
Currently, fun-asr-2025-11-07
Stable$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
fun-asr-2025-11-07
Improved far-field VAD over fun-asr-2025-08-25 for higher accuracy
Snapshot$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
fun-asr-2025-08-25Snapshot$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
fun-asr-mtl
Currently, fun-asr-mtl-2025-08-25
Stable$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
fun-asr-mtl-2025-08-25Snapshot$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
fun-asr-flash-2026-06-15
Supports synchronous calls (up to 5 minutes) and context enhancement
Snapshot$0.000035/second36,000 seconds (10 hours)
Valid for 90 days
  • Supported languages:
    • fun-asr, fun-asr-2025-11-07, fun-asr-mtl, and fun-asr-mtl-2025-08-25: 30 languages
    • fun-asr-2025-08-25: Mandarin and English.
  • Sample rates supported: Any
  • Audio formats supported: aac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv

Make your first call

Get an API key and set it as an environment variable. To use the SDK, install it.Because audio and video files are often large, file transfer and speech recognition can take a long time. The file recognition API uses asynchronous invocation to submit tasks. After the file recognition is complete, you must use the query API to retrieve the speech recognition results.

Async submit and sync wait

Submit a task and block until done.
from http import HTTPStatus
from dashscope.audio.asr import Transcription
from urllib import request
import dashscope
import os
import json

dashscope.base_http_api_url = 'https://maas.qwencloudapi.com/api/v1'

# If you have not configured an environment variable, replace the following line with your API key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.getenv("DASHSCOPE_API_KEY")

task_response = Transcription.async_call(
  model='qwen-audio-3.1-asr-flash-filetrans',
  file_urls=['https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/bjgrbu/hello_world_female_en.wav',
      'https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/rlrbee/hello_world_male_en.wav'],
  language_hints=['zh', 'en']  # language_hints is an optional parameter that specifies the language code of the audio to be recognized. For the value range, see the API reference.
)

transcription_response = Transcription.wait(task=task_response.output.task_id)

if transcription_response.status_code == HTTPStatus.OK:
  for transcription in transcription_response.output['results']:
    if transcription['subtask_status'] == 'SUCCEEDED':
      url = transcription['transcription_url']
      result = json.loads(request.urlopen(url).read().decode('utf8'))
      print(json.dumps(result, indent=4,
      ensure_ascii=False))
    else:
      print('transcription failed!')
      print(transcription)
else:
  print('Error: ', transcription_response.output.message)
First result
{
  "file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/bjgrbu/hello_world_female_en.wav",
  "properties": {
    "audio_format": "pcm_s16le",
    "channels": [
      0
    ],
    "original_sampling_rate": 24000,
    "original_duration_in_milliseconds": 3280
  },
  "transcripts": [
    {
      "channel_id": 0,
      "content_duration_in_milliseconds": 3000,
      "text": "Hello world, this is Alibaba Speech Lab. ",
      "sentences": [
        {
          "begin_time": 240,
          "end_time": 3240,
          "text": "Hello world, this is Alibaba Speech Lab. ",
          "sentence_id": 1,
          "words": [
            {
              "begin_time": 240,
              "end_time": 640,
              "text": "Hello",
              "punctuation": ""
            },
            {
              "begin_time": 640,
              "end_time": 960,
              "text": " world",
              "punctuation": ","
            },
            {
              "begin_time": 1280,
              "end_time": 1480,
              "text": " this",
              "punctuation": ""
            },
            {
              "begin_time": 1480,
              "end_time": 1840,
              "text": " is",
              "punctuation": ""
            },
            {
              "begin_time": 1840,
              "end_time": 2520,
              "text": " Alibaba",
              "punctuation": ""
            },
            {
              "begin_time": 2520,
              "end_time": 2920,
              "text": " Speech",
              "punctuation": ""
            },
            {
              "begin_time": 2920,
              "end_time": 3240,
              "text": " Lab",
              "punctuation": ". "
            }
          ]
        }
      ]
    }
  ]
}
Second result
{
  "file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/rlrbee/hello_world_male_en.wav",
  "properties": {
    "audio_format": "pcm_s16le",
    "channels": [
      0
    ],
    "original_sampling_rate": 24000,
    "original_duration_in_milliseconds": 4000
  },
  "transcripts": [
    {
      "channel_id": 0,
      "content_duration_in_milliseconds": 3160,
      "text": "Hello world, this is Alibaba Speech Lab. ",
      "sentences": [
        {
          "begin_time": 800,
          "end_time": 3960,
          "text": "Hello world, this is Alibaba Speech Lab. ",
          "sentence_id": 1,
          "words": [
            {
              "begin_time": 800,
              "end_time": 1200,
              "text": "Hello",
              "punctuation": ""
            },
            {
              "begin_time": 1200,
              "end_time": 1640,
              "text": " world",
              "punctuation": ","
            },
            {
              "begin_time": 1880,
              "end_time": 2120,
              "text": " this",
              "punctuation": ""
            },
            {
              "begin_time": 2120,
              "end_time": 2560,
              "text": " is",
              "punctuation": ""
            },
            {
              "begin_time": 2560,
              "end_time": 3360,
              "text": " Alibaba",
              "punctuation": ""
            },
            {
              "begin_time": 3360,
              "end_time": 3720,
              "text": " Speech",
              "punctuation": ""
            },
            {
              "begin_time": 3720,
              "end_time": 3960,
              "text": " Lab",
              "punctuation": ". "
            }
          ]
        }
      ]
    }
  ]
}

Async submit and async query

Submit a task and poll for results instead of blocking.
from http import HTTPStatus
from dashscope.audio.asr import Transcription
import dashscope
import os
import json

dashscope.base_http_api_url = 'https://maas.qwencloudapi.com/api/v1'

# If you have not configured an environment variable, replace the following line with your API key: dashscope.api_key = "sk-xxx"
dashscope.api_key = os.getenv("DASHSCOPE_API_KEY")

transcribe_response = Transcription.async_call(
  model='qwen-audio-3.1-asr-flash-filetrans',
  file_urls=['https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/bjgrbu/hello_world_female_en.wav',
      'https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/rlrbee/hello_world_male_en.wav']
)

while True:
  if transcribe_response.output.task_status == 'SUCCEEDED' or transcribe_response.output.task_status == 'FAILED':
    break
  transcribe_response = Transcription.fetch(task=transcribe_response.output.task_id)

if transcribe_response.status_code == HTTPStatus.OK:
  print(json.dumps(transcribe_response.output, indent=4, ensure_ascii=False))
  print('transcription done!')

RESTful API

Use any HTTP library to submit tasks and poll for results. This Python sample demonstrates the workflow:
import requests
import json
import os
import time

# If you have not configured environment variables, replace the following line with your API key: api_key = "sk-xxx"
api_key = os.getenv("DASHSCOPE_API_KEY")
file_urls = [
  "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/bjgrbu/hello_world_female_en.wav",
  "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260401/rlrbee/hello_world_male_en.wav",
]

region = "maas.qwencloudapi.com"

# Submit a file transcription task, including a list of file URLs to be transcribed
def submit_task(apikey, file_urls) -> str:

  headers = {
    "Authorization": f"Bearer {apikey}",
    "Content-Type": "application/json",
    "X-DashScope-Async": "enable",
  }
  data = {
    "model": "qwen-audio-3.1-asr-flash-filetrans",
    "input": {"file_urls": file_urls},
    "parameters": {
      "channel_id": [0],
      # "vocabulary_id": "vocab-Xxxx", # Optional, hotword ID.
    },
  }
  # URL of the audio file transcription service
  service_url = (
    f"https://{region}/api/v1/services/audio/asr/transcription"
  )
  response = requests.post(
    service_url, headers=headers, data=json.dumps(data)
  )

  # Print the response content
  if response.status_code == 200:
    return response.json()["output"]["task_id"]
  else:
    print("task failed!")
    print(response.json())
    return None


# Recursively query the task status until it is successful
def wait_for_complete(task_id):
  headers = {
    "Authorization": f"Bearer {api_key}",
    "Content-Type": "application/json",
    "X-DashScope-Async": "enable",
  }

  pending = True
  while pending:
    # URL of the task status query service
    service_url = f"https://{region}/api/v1/tasks/{task_id}"
    response = requests.post(
      service_url, headers=headers
    )
    if response.status_code == 200:
      status = response.json()['output']['task_status']
      if status == 'SUCCEEDED':
        print("task succeeded!")
        pending = False
        return response.json()['output']['results']
      elif status == 'RUNNING' or status == 'PENDING':
        pass
      else:
        print("task failed!")
        pending = False
    else:
      print("query failed!")
      pending = False
    print(response.json())
    time.sleep(0.1)


task_id = submit_task(apikey=api_key, file_urls=file_urls)
print("task_id: ", task_id)
result = wait_for_complete(task_id)
print("transcription result: ", result)

Synchronous calls (Qwen-Audio-3.x-ASR-Flash/Fun-ASR-Flash)

The Qwen-Audio-3.x-ASR-Flash and Fun-ASR-Flash model series support synchronous calls for audio files shorter than 5 minutes. Results can be returned in streaming or non-streaming mode.
curl --location --request POST 'https://maas.qwencloudapi.com/api/v1/services/aigc/multimodal-generation/generation' \
  --header "Authorization: Bearer $DASHSCOPE_API_KEY" \
  --header "Content-Type: application/json" \
  --header "X-DashScope-SSE: disable" \
  --data '{
  "model": "qwen-audio-3.1-asr-flash",
  "input": {
    "messages": [
      {
        "role": "user",
        "content": [
          {
            "type": "input_audio",
            "input_audio": {
              "data": "{YOUR_AUDIO_URL}"
            }
          }
        ]
      }
    ]
  },
  "parameters": {
    "format": "wav",
    "sample_rate": "16000"
  }
}'
Response structure note: The Qwen-Audio-3.x-ASR-Flash and Fun-ASR-Flash model series return a response structure through the DashScope synchronous API (the multimodal-generation endpoint) that differs from the standard DashScope multimodal response format. The recognized text is available at output.output.sentence.text and output.text in the response, not at output.choices[].message.content. When parsing responses from these models, use the output.output path to access the recognition results.Example response excerpt:
{
  "output": {
    "output": {
      "sentence": {
        "text": "Hello World, this is Alibaba Speech Lab."
      }
    },
    "text": "Hello World, this is Alibaba Speech Lab."
  },
  "request_id": "..."
}

Context enhancement

Supported models: Qwen-Audio-3.x-ASR-Flash (qwen-audio-3.1-asr-flash) and Fun-ASR-Flash (fun-asr-flash-2026-06-15) support context enhancement. Use case: Designed for scenarios that combine ASR with a large language model. Passing previous conversation context (LLM replies and earlier recognition results) into the ASR model significantly improves transcription accuracy for proper nouns such as names, locations, and product terms — more flexible than traditional hotwords. Usage: Pass the conversation history through input.messages. Use the assistant role for the LLM's previous replies and the user role with input_text type for earlier recognition results. Context pairs must appear before the current audio message. Supported text types include (but are not limited to):
  • Hotword lists in various delimiter formats (for example: hotword1, hotword2, hotword3, hotword4)
  • Free-form paragraphs or passages of any length
  • Mixed content: any combination of word lists and paragraphs
  • Irrelevant or meaningless text, including gibberish. The model tolerates irrelevant content well, and recognition quality rarely degrades because of it.
Example: Consider an audio clip whose correct transcription is: "What insider jargon do you know in investment banking? First, the nine top foreign investment banks, Bulge Bracket, BB..."
Without context enhancementWith context enhancement
Without context enhancement, some investment bank names are misrecognized. For example, "Bird Rock" should be "Bulge Bracket".

Result: "What insider jargon do you know in investment banking? First, the nine top foreign investment banks, Bird Rock, BB..."
With context enhancement, the investment bank names are recognized correctly.

Result: "What insider jargon do you know in investment banking? First, the nine top foreign investment banks, Bulge Bracket, BB..."
To produce the corrected result, include any of the following in the context:
  • A word list:
    • List 1:
Bulge Bracket, Boutique, Middle Market, domestic securities firms
  • List 2:
Bulge Bracket Boutique Middle Market domestic securities firms
  • List 3:
['Bulge Bracket', 'Boutique', 'Middle Market', 'domestic securities firms']
  • Natural language:
Investment bank classification: a quick guide.
Recently a few friends in Australia asked me what investment banks really are. Here is a quick primer. For students studying abroad, investment banks fall into four broad categories: Bulge Bracket, Boutique, Middle Market, and domestic securities firms.
Bulge Bracket banks: the nine top investment banks we often refer to, including Goldman Sachs, Morgan Stanley, and so on. They are large in both business scope and scale.
Boutique banks: relatively small in size but highly focused in their service areas. Firms such as Lazard and Evercore have deep expertise in specific fields.
Middle Market banks: serve mid-sized companies with M&A, IPO, and similar services. Though smaller than the bulge brackets, they hold strong positions in specific markets.
Domestic securities firms: as the Chinese market has risen, domestic firms play an increasingly important role internationally.
There are also further breakdowns by position and business line you can find in related charts. Hopefully this helps you understand investment banks and prepare for your career.
  • Natural language with distracting content: some text is unrelated to the audio, such as the names in the example below.
Investment bank classification: a quick guide.
Recently a few friends in Australia asked me what investment banks really are. Here is a quick primer. For students studying abroad, investment banks fall into four broad categories: Bulge Bracket, Boutique, Middle Market, and domestic securities firms.
Bulge Bracket banks: the nine top investment banks we often refer to, including Goldman Sachs, Morgan Stanley, and so on. They are large in both business scope and scale.
Boutique banks: relatively small in size but highly focused in their service areas. Firms such as Lazard and Evercore have deep expertise in specific fields.
Middle Market banks: serve mid-sized companies with M&A, IPO, and similar services. Though smaller than the bulge brackets, they hold strong positions in specific markets.
Domestic securities firms: as the Chinese market has risen, domestic firms play an increasingly important role internationally.
There are also further breakdowns by position and business line you can find in related charts. Hopefully this helps you understand investment banks and prepare for your career.
Wang Haoxuan, Li Zihan, Zhang Jingxing, Liu Xinyi, Chen Junjie, Yang Siyuan, Zhao Yutong, Huang Zhiqiang, Zhou Zimo, Wu Yajing, Xu Ruoxi, Sun Haoran, Hu Jinyu, Zhu Chenxi, Guo Wenbo, He Jingshu, Gao Yuhang, Lin Yifei,
Zheng Xiaoyan, Liang Bowen, Luo Jiaqi, Song Mingzhe, Xie Wanting, Tang Ziqian, Han Mengyao, Feng Yiran, Cao Qinxue, Deng Zirui, Xiao Wangshu, Xu Jiashu,
Cheng Yinuo, Yuan Zhiruo, Peng Haoyu, Dong Simiao, Fan Jingyu, Su Zijin, Lyu Wenxuan, Jiang Shihan, Ding Muchen,
Wei Shuyao, Ren Tianyou, Jiang Yichen, Hua Qingyu, Shen Xinghe, Fu Jinyu, Yao Xingchen, Zhong Lingyu, Yan Licheng, Jin Ruoshui, Tao Ranting, Qi Shaoshang, Xue Zhilan, Zou Yunfan, Xiong Ziang, Bai Wenfeng, Yi Qianfan

Advanced features

In addition to context enhancement, audio file transcription supports the following capabilities. To create and use hotwords, see Improve recognition accuracy.

Long audio file processing

Audio file transcription supports asynchronous transcription of long audio, which suits meeting minutes, interview notes, and call playback. Limits:
  • Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR/Qwen3-ASR-Flash-Filetrans: Each audio file must be no larger than 2 GB and no longer than 12 hours.
  • Qwen-Audio-3.1-ASR-Flash: Each audio file must be no larger than 2 GB and no longer than 5 minutes.
  • Qwen-Audio-3.x-ASR-Flash/Fun-ASR-Flash/Qwen3-ASR-Flash: Each audio file must be no larger than 10 MB and no longer than 5 minutes. For longer audio, use Qwen-Audio-3.1-ASR-Flash-Filetrans, Fun-ASR, or Qwen3-ASR-Flash-Filetrans.
  • When speaker diarization is enabled: Audio longer than 2 hours is not recommended, because recognition may fail or time out. See Speaker diarization.
How it works: Long audio transcription uses an asynchronous task model with three steps:
  1. Submit the transcription task and get a task_id.
  2. Poll for the task status, or use the SDK's wait method to block until the task finishes.
  3. When the task completes, download the result JSON from the returned URL.
For code samples, see Getting started. For polling intervals and batch task management, see Asynchronous task management.

Speaker diarization

Speaker diarization identifies different speakers in the audio and labels each sentence in the transcription result, which suits multi-person meetings and interview recordings. Qwen-Audio-3.1-ASR-Flash enables speaker diarization with speaker_diarization_enabled. For its parameters and response structure, see RESTful API. Supported models: Qwen-Audio-3.x-ASR-Flash-Filetrans, Fun-ASR, and Paraformer. How to enable: Set diarization_enabled to true in the request parameters. Each sentence in the result then includes a speaker_id field that identifies the speaker. Response structure (excerpt):
{
  "transcripts": [
    {
      "sentences": [
        { "begin_time": 100, "end_time": 3820, "text": "Hello, let's review the project progress today.", "speaker_id": 0 },
        { "begin_time": 3820, "end_time": 6500, "text": "Sure, I'll start with my update.", "speaker_id": 1 }
      ]
    }
  ]
}
When speaker diarization is enabled, audio longer than 2 hours is not recommended, because recognition may fail or time out. For the duration limits that apply when diarization is off, see Long audio file processing. Speaker diarization supports mono audio only.
SDKs expose these fields with different naming conventions (dictionary keys, object properties, or methods). For the full field mapping, see API reference.

Sensitive word filtering

Sensitive word filtering replaces or removes sensitive words in the recognition result, which suits customer service quality checks, content compliance, and subtitle review. Supported models: Qwen-Audio-3.x-ASR-Flash-Filetrans, Fun-ASR, and Paraformer. Default behavior: If you do not pass the special_word_filter parameter, the built-in sensitive word list applies and matched words are replaced with an equal number of asterisks (*). Custom configuration: special_word_filter is a JSON object with three sub-fields:
  • filter_with_signed.word_list: A string array of sensitive words to replace with an equal number of asterisks. For example, with ["test"], "Run a test for me" becomes "Run a **** for me".
  • filter_with_empty.word_list: A string array of sensitive words to remove entirely from the result. For example, with ["start"], "Is the game about to start" becomes "Is the game about to".
  • system_reserved_filter: A boolean that defaults to true. Specifies whether to also apply the built-in sensitive word list, which is combined with your custom lists.
Example configuration:
{
  "special_word_filter": {
    "filter_with_signed": {
      "word_list": ["test"]
    },
    "filter_with_empty": {
      "word_list": ["start", "happen"]
    },
    "system_reserved_filter": true
  }
}
The default differs from real-time speech recognition: in real-time recognition, no filtering is applied when special_word_filter is omitted, because system_reserved_filter defaults to false.

Emotion recognition

Qwen3-ASR-Flash-Filetrans and Qwen3-ASR-Flash always perform emotion recognition, with no extra configuration. The result includes the speaker's emotion label, which is one of seven values: surprised, neutral, happy, sad, disgusted, angry, and fearful. Field paths (they vary by interface):
  • OpenAI-compatible interface (Qwen3-ASR-Flash): choices[].delta.annotations[].emotion for streaming output, or choices[].message.annotations[].emotion for non-streaming output.
  • DashScope synchronous interface (Qwen3-ASR-Flash): output.choices[].message.annotations[].emotion.
  • DashScope asynchronous task interface (Qwen3-ASR-Flash-Filetrans): transcripts[].sentences[].emotion, alongside the timestamp fields in each sentence object.
Response structure (DashScope asynchronous task interface, excerpt):
{
  "transcripts": [{
    "sentences": [{
      "begin_time": 0,
      "end_time": 1440,
      "text": "Welcome to QwenCloud.",
      "emotion": "neutral",
      "language": "en"
    }]
  }]
}
Qwen-Audio-3.x-ASR-Flash-Filetrans, Qwen-Audio-3.x-ASR-Flash, Fun-ASR-Flash, Fun-ASR, and Paraformer do not support emotion recognition. To use emotion recognition in real-time scenarios, see Real-time speech recognition.

Get timestamps

Audio file transcription can return timestamps in the result, which suits subtitle generation, keyword highlighting, and audio or video editing. The default behavior and controls vary by model:
  • Qwen-Audio-3.x-ASR-Flash-Filetrans/Qwen-Audio-3.x-ASR-Flash/Fun-ASR/Fun-ASR-Flash: Timestamps are always on and cannot be turned off.
  • Qwen3-ASR-Flash-Filetrans: Timestamps are supported only through the DashScope asynchronous interface, where they are always on. Use the enable_words request parameter to control the granularity: false (default) returns sentence-level timestamps, and true returns word-level timestamps. Word-level timestamps are supported only for Chinese, English, Japanese, Korean, German, French, Spanish, Italian, Portuguese, and Russian. Accuracy is not guaranteed for other languages.
When you call Qwen3-ASR-Flash through the OpenAI-compatible interface, the output is a chat.completion object and contains no timestamp fields. If you need timestamps, use Qwen3-ASR-Flash-Filetrans through the asynchronous task interface.
Timestamps are in milliseconds and are returned at two levels:
  • Sentence level: sentences[].begin_time and sentences[].end_time mark the start and end of each sentence in the audio.
  • Word level: The sentences[].words[] array, where each element contains begin_time, end_time, and text (the word).
Response structure (DashScope asynchronous task interface, excerpt):
{
  "transcripts": [{
    "sentences": [{
      "begin_time": 100,
      "end_time": 3820,
      "text": "Hello, let's review the project progress today.",
      "words": [
        { "begin_time": 100, "end_time": 596, "text": "Hello" },
        { "begin_time": 596, "end_time": 844, "text": "let's" }
      ]
    }]
  }]
}
Timestamps inside the audio are millisecond integers such as 100. Do not confuse them with the task-level end_time, which is the task completion time as a date string such as "2024-09-12 15:11:40.903".

Going live

When you move audio file transcription to production, the following practices improve recognition quality and system stability.
  • File hosting: Upload audio files to an object storage service such as OSS and call the API by URL. Avoid local file uploads, which are capped at 100 QPS and cannot be scaled up.
  • Asynchronous polling: Long audio transcription is asynchronous, and the task query API is limited to 20 QPS by default. Set a reasonable polling interval, such as 2 to 5 seconds, to avoid triggering rate limits. For polling strategies and batch task management, see Asynchronous task management.
  • Error handling: Implement a robust retry mechanism. For network timeouts or temporary server-side errors (5xx), retry with an exponential backoff strategy.
  • Noise reduction: For noisy audio, preprocess it with a tool such as FFmpeg before you submit it for recognition.
  • Model selection: Choose the model based on audio duration. For audio within 5 minutes, use Qwen3-ASR-Flash. For audio longer than 5 minutes, use Qwen-Audio-3.1-ASR-Flash-Filetrans, Fun-ASR, or Qwen3-ASR-Flash-Filetrans.

Compare models

FeatureFun-ASR
Supported languagesVaries by model: fun-asr, fun-asr-2025-11-07, fun-asr-mtl, fun-asr-mtl-2025-08-25: Chinese (Mandarin, Cantonese, Wu, Minnan, Hakka, Gan, Xiang, Jin; also supports accents from Central Plains, Southwest, Ji-Lu, Jianghuai, Lan-Yin, Jiao-Liao, Northeast, Beijing, Hong Kong, and Taiwan, including official dialects from regions such as Henan, Shaanxi, Hubei, Sichuan, Chongqing, Yunnan, Guizhou, Guangdong, Guangxi, Hebei, Tianjin, Shandong, Anhui, Nanjing, Jiangsu, Hangzhou, Gansu, and Ningxia), English, Japanese, Korean, Vietnamese, Indonesian, Thai, Malay, Filipino, Arabic, Bulgarian, Croatian, Czech, Danish, Dutch, Estonian, Finnish, Greek, Hindi, Hungarian, Irish, Latvian, Lithuanian, Maltese, Polish, Portuguese, Romanian, Slovak, Slovenian, Swedish. fun-asr-2025-08-25: Chinese (Mandarin), English
Supported audio formatsaac, amr, avi, flac, flv, m4a, mkv, mov, mp3, mp4, mpeg, ogg, opus, wav, webm, wma, wmv
Sample rateAny
Sound channelsAny
Input formatPublicly accessible URLs of files to be recognized. Up to 100 audio files are supported.
Audio size/durationEach audio file must be no larger than 2 GB and no longer than 12 hours.
Emotion recognitionNot supported
TimestampSupported (always on)
Punctuation predictionSupported (always on)
HotwordsSupported. The hotword feature is supported only in the primary workspace and is not available in sub-workspaces.
ITNSupported (always on)
Singing voice recognitionSupported (fun-asr and fun-asr-2025-11-07 only)
Noise rejectionSupported (always on)
Sensitive word filteringSupported (filters content from the QwenCloud sensitive word list by default)
Speaker diarizationSupported (off by default, can be enabled)
Filler word filteringNot supported
VADSupported (always on)
Rate limiting (RPS)Job submission API: 10, Task query API: 20
Connection typesDashScope: Java/Python SDK, RESTful API
PricingInternational: $0.000035/second

API reference

FAQ

  • Qwen-Audio-3.x-ASR-Flash-Filetrans/Fun-ASR
  • Qwen-ASR
  • Qwen-Omni

How can I improve recognition accuracy?

Several factors affect accuracy. Review each and apply the corresponding optimization.Key factors:
  1. Sound quality: Recording device quality, sample rate, and ambient noise directly affect clarity. High-quality audio input is essential.
  2. Speaker characteristics: Variations in pitch, speech rate, accent, and dialect increase recognition difficulty, especially for rare dialects or heavy accents.
  3. Language and vocabulary: Mixed languages, technical terms, or slang increase recognition difficulty. Configure hotwords to improve accuracy for domain-specific terms.
  4. Contextual understanding: Insufficient context can cause semantic ambiguity, especially in situations where surrounding context is needed for correct recognition.
Optimization methods:
  1. Optimize audio quality: Use high-performance microphones at the recommended sample rate. Minimize ambient noise and echo.
  2. Adapt to the speaker: For audio with strong accents or dialects, select a model that supports those specific dialects.
  3. Configure hotwords: Set hotwords for technical terms, proper nouns, and other specific words. For more information, see Customize hotwords.
  4. Preserve context: Avoid splitting audio into excessively short clips.