# QwenCloud Docs > QwenCloud provides API access to Qwen large language models and multimodal AI. Capabilities include text generation, vision, image generation, video generation, speech-to-text, text-to-speech, speech-to-speech, embeddings, and reranking via OpenAI-compatible, Anthropic-compatible, and DashScope endpoints. ## Token Plan - [Token Plan](https://docs.qwencloud.com/token-plan/overview.md): Token Plan is an AI model subscription service from QwenCloud, using Credits as a unified billing unit across text, image, and video generation models. Available in Personal Edition and Team Edition. - [Token Plan Individual](https://docs.qwencloud.com/token-plan/personal/token-plan-personal-overview.md): Token Plan Individual is an AI model subscription service for individual developers with Credits-based billing across text, multimodal models and Harness tools, compatible with mainstream AI programming and agent tools. - [Quick start](https://docs.qwencloud.com/token-plan/personal/token-plan-personal-quickstart.md): Get started with Token Plan Individual in three steps. - [FAQ](https://docs.qwencloud.com/token-plan/personal/token-plan-personal-faq.md): Frequently asked questions about Token Plan Individual. - [Token Plan Team Edition](https://docs.qwencloud.com/token-plan/team/token-plan-team-overview.md): AI model subscription service for teams with Credits-based billing across text, image, and video generation models - [Quick start](https://docs.qwencloud.com/token-plan/team/token-plan-team-quickstart.md): Start using Token Plan (Team Edition) in three steps: subscribe, get your API key, and set up your AI tool. - [Team management](https://docs.qwencloud.com/token-plan/team/team-management.md): Add and manage team members, assign and revoke seats, configure SSO login, and monitor Credits usage in the Token Plan Management Console. - [FAQ](https://docs.qwencloud.com/token-plan/team/token-plan-team-faq.md): Fix Token Plan issues fast ## Best Practices - [Integrate Harness tools](https://docs.qwencloud.com/token-plan/best-practices/built-in-tools.md): Extend AI coding tools with built-in Harness tools in supported Qwen models - [Integrate multimodal generation models](https://docs.qwencloud.com/token-plan/best-practices/integrate-multimodal-gen.md): Integrate image generation, video generation, and speech synthesis models with Token Plan through tool extension mechanisms - [Add visual understanding capabilities](https://docs.qwencloud.com/token-plan/best-practices/add-vision-skill.md): Vision for coding models ## Getting Started - [Build with QwenCloud](https://docs.qwencloud.com/developer-guides/getting-started/introduction.md): AI models for text, vision, speech, and image & video generation. - [First API call](https://docs.qwencloud.com/developer-guides/getting-started/first-api-call.md): Get started in a few minutes - [Choose models](https://docs.qwencloud.com/developer-guides/getting-started/model-selection.md): Match your specific use case - [Pricing](https://docs.qwencloud.com/developer-guides/getting-started/pricing.md): Pay-as-you-go pricing for API usage - [Latest model: Qwen3.8-Flash](https://docs.qwencloud.com/developer-guides/getting-started/latest-model.md): Learn about the Qwen3.8-Flash model — its next-gen architecture, capabilities, specs, and how to get started ## Models - [Text generation models](https://docs.qwencloud.com/developer-guides/getting-started/text-generation-models.md): Choose a model for AI agents, chatbots, document processing, and more. - [Generate text](https://docs.qwencloud.com/developer-guides/text-generation/quickstart.md): Make your first text generation call - [Partial mode](https://docs.qwencloud.com/developer-guides/text-generation/partial-mode.md): Continue from a prefix - [Machine translation (Qwen-MT)](https://docs.qwencloud.com/developer-guides/text-generation/qwen-mt.md): 92 languages with terms - [Role playing (Qwen-Character)](https://docs.qwencloud.com/developer-guides/text-generation/role-playing.md): NPCs and virtual personas - [DeepSeek](https://docs.qwencloud.com/developer-guides/third-party-models/deepseek.md): Call DeepSeek models through the OpenAI-compatible API or DashScope SDK on QwenCloud. - [GLM](https://docs.qwencloud.com/developer-guides/third-party-models/glm.md): Call GLM models through the OpenAI-compatible API or DashScope SDK on QwenCloud. - [GLM-ZHIPU](https://docs.qwencloud.com/developer-guides/third-party-models/glm-zhipu.md): This document describes how to call the Z.AI GLM model inference service on QwenCloud. - [Kimi](https://docs.qwencloud.com/developer-guides/third-party-models/kimi.md): Call the Kimi K3 and Kimi K2.7 Code models through the OpenAI-compatible API or DashScope SDK on QwenCloud. - [Visual understanding models](https://docs.qwencloud.com/developer-guides/getting-started/vision-models.md): Choose a model for image analysis, video understanding, OCR, and more. - [Analyze images and videos](https://docs.qwencloud.com/developer-guides/multimodal/vision.md): Generate content from visual inputs - [Text extraction](https://docs.qwencloud.com/developer-guides/multimodal/ocr.md): OCR for docs and tables - [Image generation models](https://docs.qwencloud.com/developer-guides/getting-started/image-models.md): Choose a model for text-to-image generation, image editing, and more. - [Text-to-image](https://docs.qwencloud.com/developer-guides/image-generation/text-to-image.md): Generate images from text prompts. - [Image editing](https://docs.qwencloud.com/developer-guides/image-generation/image-editing.md): Modify images via text - [Image editing - Wan2.7/2.6/2.5](https://docs.qwencloud.com/developer-guides/image-generation/wan-image-editing.md): Edit images using text instructions with Wan2.7, 2.6, and 2.5 models. - [Video generation models](https://docs.qwencloud.com/developer-guides/getting-started/video-models.md): Choose a model for text-to-video, image-to-video, reference-to-video, video editing, and more. - [Wan3.0 Video Generation](https://docs.qwencloud.com/developer-guides/video-generation/wan30-video.md): The wan3.0 series is an All-in-One video generation model supporting text-to-video, image-to-video, multi-modal reference, video editing, and video extension. Up to 30 seconds per generation, 30fps, with native audio output. - [Text-to-video](https://docs.qwencloud.com/developer-guides/video-generation/text-to-video.md): Generate video from text - [Image-to-video: first frame](https://docs.qwencloud.com/developer-guides/video-generation/image-to-video.md): Animate from a single image - [Image-to-video: first and last frames](https://docs.qwencloud.com/developer-guides/video-generation/image-to-video-first-last.md): Animate between two frames - [Reference-to-video](https://docs.qwencloud.com/developer-guides/video-generation/reference-video.md): Replicate motion and look - [General video editing](https://docs.qwencloud.com/developer-guides/video-generation/video-editing.md): Repaint, extend, and edit - [Text-to-speech models](https://docs.qwencloud.com/developer-guides/speech/tts-models.md): Choose a model for speech synthesis, voice cloning, and voice design. - [Real-time speech synthesis](https://docs.qwencloud.com/developer-guides/speech/realtime-streaming.md): Stream text-to-speech conversion with low first-packet latency - [Non-real-time speech synthesis](https://docs.qwencloud.com/developer-guides/speech/tts.md): Non-realtime speech synthesis with Qwen3-TTS - [Voice cloning](https://docs.qwencloud.com/developer-guides/speech/voice-cloning.md): Clone a voice from audio samples for use with Qwen-Audio-TTS, CosyVoice, Qwen-TTS, or Qwen-Omni models. - [Voice design](https://docs.qwencloud.com/developer-guides/speech/voice-design.md): Create custom voices from text descriptions for use with CosyVoice or Qwen TTS models. - [SSML and LaTeX](https://docs.qwencloud.com/developer-guides/speech/ssml.md): Control speech rate, pauses, and pronunciation with SSML, or convert LaTeX formulas to natural speech - [Qwen-Audio-TTS voice list](https://docs.qwencloud.com/developer-guides/speech/voice-list/qwen-audio-tts.md): System and base voices for qwen-audio-3.0-tts-plus and qwen-audio-3.0-tts-flash - [CosyVoice voice list](https://docs.qwencloud.com/developer-guides/speech/voice-list/cosyvoice.md): System voice catalog - [Qwen-TTS voice list](https://docs.qwencloud.com/developer-guides/speech/voice-list/qwen-tts.md): Qwen-TTS voice reference for real-time and non-real-time speech synthesis - [Speech-to-text models](https://docs.qwencloud.com/developer-guides/speech/speech-to-text-models.md): Choose a model for live captions, file transcription, and more. - [Realtime speech recognition](https://docs.qwencloud.com/developer-guides/speech/asr-realtime.md): Live speech to text - [Audio file transcription](https://docs.qwencloud.com/developer-guides/speech/asr.md): Convert files to text - [Improve recognition accuracy](https://docs.qwencloud.com/developer-guides/speech/improve-recognition-accuracy.md): Improve speech recognition accuracy using custom hotwords and context enhancement. - [Speech-to-speech models](https://docs.qwencloud.com/developer-guides/speech/s2s-models.md): Choose a model for voice conversation, speech translation, or simultaneous interpretation. - [Realtime Audio Chat (Qwen-Audio-Realtime)](https://docs.qwencloud.com/developer-guides/speech/qwen-audio-realtime.md): Qwen-Audio is an end-to-end real-time voice interaction model for low-latency voice conversations. Use cases include voice assistants, intelligent customer service, and AI companions. - [Real-time audio and video translation](https://docs.qwencloud.com/developer-guides/speech/realtime-translation.md): Real-time speech translation with 2.8-second latency - [Audio and video file translation](https://docs.qwencloud.com/developer-guides/speech/file-translation.md): 18-language translation - [Omni-modal models](https://docs.qwencloud.com/developer-guides/speech/omni-models.md): Choose a model for multimodal understanding, audio and video analysis, voice conversation, content moderation, or speech translation. - [Realtime audio and video understanding](https://docs.qwencloud.com/developer-guides/speech/realtime-multimodal-speech.md): Live audio and video chat - [Audio and video file understanding](https://docs.qwencloud.com/developer-guides/speech/multimodal-speech.md): Text + image/audio input - [Omni-modal voice list](https://docs.qwencloud.com/developer-guides/speech/omni-voice-list.md): Voices supported by Qwen3.5-Omni, Qwen3-Omni-Flash, Livetranslate, and other omni-modal models, with their voice parameter values. - [Embedding & reranking models](https://docs.qwencloud.com/developer-guides/getting-started/embedding-models.md): Choose a model for semantic search, RAG retrieval, cross-modal matching, and reranking. - [Text and multimodal embedding](https://docs.qwencloud.com/developer-guides/embeddings/embedding.md): Embedding models convert text, images, and video into numerical vectors for semantic search, recommendations, clustering, classification, and anomaly detection. - [Reranking](https://docs.qwencloud.com/developer-guides/embeddings/reranking.md): Improve search accuracy - [Function Calling](https://docs.qwencloud.com/developer-guides/tool-calling/function-calling.md): Connect models to external tools - [Web search](https://docs.qwencloud.com/developer-guides/tool-calling/web-search.md): Ground model responses in real-time web data - [Web extractor](https://docs.qwencloud.com/developer-guides/tool-calling/web-scraping.md): Fetch URLs for context - [Code interpreter](https://docs.qwencloud.com/developer-guides/tool-calling/code-interpreter.md): Run Python in a sandbox - [Image search](https://docs.qwencloud.com/developer-guides/tool-calling/image-search.md): Find images via Responses - [Connect to MCP servers](https://docs.qwencloud.com/developer-guides/tool-calling/mcp.md): Use external tools in chat - [Structured output](https://docs.qwencloud.com/developer-guides/text-generation/structured-output.md): Make the model return valid JSON, and use JSON Schema to constrain the output structure precisely - [Thinking](https://docs.qwencloud.com/developer-guides/text-generation/thinking.md): Solve complex tasks with step-by-step thinking - [Batch API](https://docs.qwencloud.com/developer-guides/text-generation/batch.md): Process bulk requests asynchronously at 50% off ## Run and Scale - [Prime Mode](https://docs.qwencloud.com/developer-guides/run-and-scale/prime-mode.md): Prime Mode delivers higher TPS for latency-sensitive scenarios such as AI coding assistants, multi-step agent reasoning, and real-time conversations. - [Multi-turn conversations](https://docs.qwencloud.com/developer-guides/run-and-scale/multi-turn.md): Manage chat context - [Context cache](https://docs.qwencloud.com/developer-guides/run-and-scale/context-cache.md): Cut cost with prefix reuse - [Streaming output](https://docs.qwencloud.com/developer-guides/run-and-scale/streaming.md): Receive text output token by token as it is generated. - [Async task management](https://docs.qwencloud.com/developer-guides/run-and-scale/async-task-management.md): Two async patterns on QwenCloud: task-based for media generation, batch for high-volume text processing - [Connection reuse and pooling](https://docs.qwencloud.com/developer-guides/run-and-scale/connection-pooling.md): HTTP connection reuse and WebSocket connection pooling for high-concurrency workloads. ## Tutorials - [Accuracy tuning](https://docs.qwencloud.com/developer-guides/accuracy-tuning/overview.md): Maximize correctness and consistent behavior across text, image, video, speech, and vision models on QwenCloud - [Text-to-text](https://docs.qwencloud.com/developer-guides/accuracy-tuning/text-generation.md): Design effective prompts - [Text-to-image](https://docs.qwencloud.com/developer-guides/accuracy-tuning/image-generation.md): Better Wan image results - [Text-to-video](https://docs.qwencloud.com/developer-guides/accuracy-tuning/video-generation.md): Craft better video prompts - [Speech recognition best practices](https://docs.qwencloud.com/developer-guides/accuracy-tuning/speech-recognition.md): Optimize ASR quality through audio best practices, vocabulary customization, and post-processing correction. - [Explicit Cache Best Practices](https://docs.qwencloud.com/developer-guides/accuracy-tuning/explicit-cache-best-practice.md): Explicit cache guarantees deterministic cache hits for identical input content by adding cache markers to your requests, significantly reducing cost and latency. - [Latency optimization](https://docs.qwencloud.com/developer-guides/run-and-scale/latency-optimization.md): Faster response times across text, image, video, and speech models on QwenCloud - [Cost optimization](https://docs.qwencloud.com/developer-guides/run-and-scale/cost-optimization.md): Spend less on text, image, video, and speech API calls while maintaining output quality - [Safety](https://docs.qwencloud.com/developer-guides/run-and-scale/safety.md): Content moderation, input/output guardrails, and responsible AI practices across all modalities - [Realtime call with qwen3.5-omni-plus-realtime via WebRTC](https://docs.qwencloud.com/developer-guides/tutorials/realtime/webrtc-omni-realtime.md): Use qwen3.5-omni-plus-realtime model via WebRTC protocol for real-time calls - [Realtime call with qwen3.5-omni-plus-realtime via AOQ](https://docs.qwencloud.com/developer-guides/tutorials/realtime/aoq-omni-realtime.md): Use qwen3.5-omni-plus-realtime model via AOQ protocol for real-time calls - [Build push-to-talk voice conversations with qwen3.5-omni-plus-realtime over AOQ](https://docs.qwencloud.com/developer-guides/tutorials/realtime/aoq-omni-ptt.md): Use AOQ to connect to qwen3.5-omni-plus-realtime and let the client control turn boundaries for push-to-talk conversations and optional image questions. The client code uses iOS Swift. - [Real-time speech recognition with AOQ + fun-asr-realtime](https://docs.qwencloud.com/developer-guides/tutorials/realtime/aoq-fun-asr-realtime.md): Use AOQ to connect to fun-asr-realtime for streaming microphone audio and receiving real-time speech recognition results. Client code examples use Android Java; other AOQ-supported platforms share the same interface. - [Real-time voice conversation with AOQ + Qwen-Audio](https://docs.qwencloud.com/developer-guides/tutorials/realtime/aoq-audio-realtime.md): Use AOQ to connect to qwen-audio-3.0-realtime-plus and use server-side VAD to build low-latency real-time voice conversations. Client code examples use Android Java. - [Synthesize speech with qwen-audio-3.0-tts-flash over AOQ](https://docs.qwencloud.com/developer-guides/tutorials/realtime/aoq-tts-realtime.md): Use AOQ to connect to qwen-audio-3.0-tts-flash, send text in segments, and play synthesized speech in real time. The client code uses Android Java. ## Administration - [API keys](https://docs.qwencloud.com/developer-guides/administration/api-keys.md): Create and manage API keys - [Workspaces](https://docs.qwencloud.com/developer-guides/administration/workspace.md): Organize users and access - [Rate limits](https://docs.qwencloud.com/developer-guides/administration/rate-limits.md): Understand and manage API rate limits ## Integrations - [OpenClaw](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/openclaw.md): Open-source AI assistant platform - [Hermes Agent](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/hermes-agent.md): Terminal AI coding assistant by Nous Research - [Claude Code](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/claude-code.md): Anthropic's terminal AI coding assistant - [OpenCode](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/opencode.md): Terminal AI coding assistant - [Cursor](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/cursor.md): AI-powered code editor - [Codex](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/codex.md): OpenAI's terminal AI coding assistant - [Qwen Code](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/qwen-code.md): Official terminal-based AI coding tool - [DeepSeek Harness](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/deepseek-harness.md): DeepSeek Harness is an open-source AI Agent framework by DeepSeek, supporting both Web UI and CLI modes. Connect it to QwenCloud using Pay-as-you-go, Token Plan Personal Edition, or Token Plan Team Edition. - [QwenPaw](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/qwenpaw.md): Open-source personal AI assistant from the AgentScope team - [Cherry Studio](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/cherry-studio.md): Open-source AI desktop client - [Chatbox](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/chatbox.md): Cross-platform AI chat client - [Cline](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/cline.md): VSCode AI coding extension - [Qoder](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/qoder.md): Agentic coding platform with IDE, CLI, and JetBrains plugin - [Qoder CN (formerly Lingma)](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/lingma.md): Alibaba Cloud intelligent coding assistant IDE - [Kilo CLI](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/kilo-cli.md): AI coding in terminal - [Postman](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/postman.md): API testing tool - [Dify](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/dify.md): Low-code LLM app platform - [More tools](https://docs.qwencloud.com/developer-guides/clients-and-developer-tools/other-tools.md): Connect any OpenAI or Anthropic compatible programming tool - [MLOps & observability](https://docs.qwencloud.com/developer-guides/integrations/mlops-observability.md): Production AI monitoring - [Monitoring alerts](https://docs.qwencloud.com/developer-guides/integrations/alerts.md): Configure alert rules for model API metrics ## Model Production - [Datasets overview](https://docs.qwencloud.com/developer-guides/datasets/overview.md): Manage training datasets for fine-tuning models on QwenCloud. - [Create a dataset](https://docs.qwencloud.com/developer-guides/datasets/create-dataset.md): Upload training data to QwenCloud for use in fine-tuning jobs. - [Manage datasets](https://docs.qwencloud.com/developer-guides/datasets/manage-datasets.md): Publish, edit, and delete datasets on the QwenCloud platform. - [Fine-tuning overview](https://docs.qwencloud.com/developer-guides/fine-tuning/overview.md): Improve model performance for your use case by fine-tuning Qwen models on QwenCloud with SFT and LoRA. - [Create a fine-tuning job](https://docs.qwencloud.com/developer-guides/fine-tuning/create-fine-tuning-job.md): Step-by-step guide to creating a fine-tuning job in the QwenCloud console. - [Manage fine-tuning jobs](https://docs.qwencloud.com/developer-guides/fine-tuning/manage-fine-tuning-jobs.md): Monitor, manage, and deploy fine-tuning jobs from the QwenCloud console. - [Hyperparameters reference](https://docs.qwencloud.com/developer-guides/fine-tuning/hyperparameters.md): Reference for all fine-tuning hyperparameters available in Expert mode on QwenCloud. - [Custom models overview](https://docs.qwencloud.com/developer-guides/custom-models/overview.md): View and manage models you have fine-tuned on QwenCloud. - [Model evaluation overview](https://docs.qwencloud.com/developer-guides/evaluation/overview.md): Evaluate model performance across multiple dimensions on QwenCloud, with LLM-based scoring and manual annotation. - [Evaluation dimensions](https://docs.qwencloud.com/developer-guides/evaluation/evaluation-dimensions.md): Create and manage evaluation dimensions on QwenCloud with LLM numeric scoring, LLM classification, and manual annotation. - [Evaluation tasks](https://docs.qwencloud.com/developer-guides/evaluation/evaluation-tasks.md): Create evaluation tasks on QwenCloud to quantify model output quality using custom dimensions and datasets. - [Deployment overview](https://docs.qwencloud.com/developer-guides/deployment/overview.md): Deploy custom models on QwenCloud to create dedicated inference services for production workloads. - [Manage deployments](https://docs.qwencloud.com/developer-guides/deployment/manage-deployments.md): Monitor, scale, and manage the lifecycle of your model deployments on QwenCloud. ## Security & Compliance - [Data security & privacy](https://docs.qwencloud.com/developer-guides/security-compliance/data-security.md): How QwenCloud secures your data - [Audit & access logs](https://docs.qwencloud.com/developer-guides/security-compliance/audit-logs.md): Track API usage for compliance ## Getting Started - [Get an API key](https://docs.qwencloud.com/api-reference/preparation/api-key.md): First step to using models - [Configure your API key](https://docs.qwencloud.com/api-reference/preparation/export-api-key-env.md): Avoid hardcoding secrets - [Install the SDK](https://docs.qwencloud.com/api-reference/preparation/install-sdk.md): Python and Java setup - [CLI Tool](https://docs.qwencloud.com/api-reference/preparation/cli.md): QwenCloud CLI for management and model invocation: manage the model catalog, accounts, usage, billing, subscriptions, and support tickets, and invoke text, image, video, and speech models - [Error messages](https://docs.qwencloud.com/api-reference/preparation/error-messages.md): API error code reference ## Chat Models - [OpenAI chat](https://docs.qwencloud.com/api-reference/chat/openai-chat.md): Compatible Chat API - [OpenAI responses](https://docs.qwencloud.com/api-reference/chat/openai-responses.md): Compatible Responses API - [Anthropic Messages API](https://docs.qwencloud.com/api-reference/chat/anthropic.md): Call Qwen models using Anthropic SDKs - [DashScope chat](https://docs.qwencloud.com/api-reference/chat/dashscope.md): Native SDK and HTTP API ## Image Generation API - [Qwen synchronous](https://docs.qwencloud.com/api-reference/image-generation/qwen-text-to-image.md): Real-time image generation - [Qwen — Async image generation (3.0)](https://docs.qwencloud.com/api-reference/image-generation/qwen-text-to-image-30-async.md): Asynchronous image generation for qwen-image-3.0 series - [Qwen — Generate an image](https://docs.qwencloud.com/api-reference/image-generation/qwen-text-to-image-async.md): Async image generation - [Qwen — Retrieve image result](https://docs.qwencloud.com/api-reference/image-generation/qwen-text-to-image-task-query.md): Check image task status - [Z-Image](https://docs.qwencloud.com/api-reference/image-generation/z-image.md): Fast lightweight generation - [Wan T2I synchronous](https://docs.qwencloud.com/api-reference/image-generation/wan-text-to-image-v2/synchronous.md): Real-time Wan text-to-image generation - [Wan v2 — Generate an image](https://docs.qwencloud.com/api-reference/image-generation/wan-text-to-image-v2/create-task.md): Async Wan image generation - [Wan v2 — Retrieve image result](https://docs.qwencloud.com/api-reference/image-generation/wan-text-to-image-v2/query-result.md): Check Wan image status - [Qwen](https://docs.qwencloud.com/api-reference/image-generation/qwen-image-editing.md): Edit images via text - [Wan 2.7 synchronous](https://docs.qwencloud.com/api-reference/image-generation/wan27-image-gen-edit/synchronous.md): Synchronous Wan 2.7 image generation and editing - [Wan 2.7 — Generate or edit an image](https://docs.qwencloud.com/api-reference/image-generation/wan27-image-gen-edit/create-task.md): Async Wan 2.7 image generation and editing - [Wan 2.7 — Retrieve image result](https://docs.qwencloud.com/api-reference/image-generation/wan27-image-gen-edit/query-result.md): Check Wan 2.7 image task status - [Wan 2.6 synchronous](https://docs.qwencloud.com/api-reference/image-generation/wan26-image-gen-edit/synchronous.md): Synchronous Wan 2.6 image generation and editing - [Wan 2.6 — Generate or edit an image](https://docs.qwencloud.com/api-reference/image-generation/wan26-image-gen-edit/create-task.md): Async Wan 2.6 image generation and editing - [Wan 2.6 — Retrieve image result](https://docs.qwencloud.com/api-reference/image-generation/wan26-image-gen-edit/query-result.md): Check Wan 2.6 image task status - [Wan 2.5 — Edit an image](https://docs.qwencloud.com/api-reference/image-generation/wan25-general-image-editing/create-task.md): Async Wan2.5 image edit - [Wan 2.5 — Retrieve editing result](https://docs.qwencloud.com/api-reference/image-generation/wan25-general-image-editing/query-result.md): Check Wan2.5 edit status ## Image Translation - [Qwen-MT-Image — Synchronous](https://docs.qwencloud.com/api-reference/image-translation/qwen-mt-image/synchronous.md): Call qwen-mt-image-2.0 image translation synchronously and receive the translated image URL directly - [Create image translation task](https://docs.qwencloud.com/api-reference/image-translation/qwen-mt-image/create-task.md): Submit a qwen-mt-image async image translation task - [Query task result](https://docs.qwencloud.com/api-reference/image-translation/qwen-mt-image/query-result.md): Query qwen-mt-image image translation task status and result ## Video Generation API - [wan3.0-video Video Generation](https://docs.qwencloud.com/api-reference/video-generation/wan30-video/create-task.md): wan3.0-video All-in-One video generation model. Supports text-to-video, first-frame/first+last-frame, all-in-one reference, audio, and more. - [wan3.0-video Query Task Result](https://docs.qwencloud.com/api-reference/video-generation/wan30-video/query-result.md): Query the status and result of a wan3.0-video video generation task. - [HappyHorse -- Generate a video from an image](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-image-to-video/create-task.md): Submit HappyHorse image-to-video task (first frame) - [HappyHorse -- Retrieve image-to-video result](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-image-to-video/query-result.md): Check HappyHorse image-to-video task status - [Wan 2.7 -- Generate a video from image](https://docs.qwencloud.com/api-reference/video-generation/wan27-image-to-video/create-task.md): Submit image-to-video task (wan2.7) - [Wan 2.7 -- Retrieve image-to-video result](https://docs.qwencloud.com/api-reference/video-generation/wan27-image-to-video/query-result.md): Check wan2.7 image-to-video task status - [Wan — Animate from first frame](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-video-first-frame/create-task.md): Submit image-to-video task - [Wan — Retrieve first-frame result](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-video-first-frame/query-result.md): Check video task status - [Wan — Animate from first & last frames](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-video-first-last-frames/create-task.md): Submit first+last frames - [Wan — Retrieve first-last-frame result](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-video-first-last-frames/query-result.md): Check video task status - [HappyHorse -- Generate a video from reference images](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-reference-to-video/create-task.md): Submit HappyHorse reference-to-video task - [HappyHorse -- Retrieve reference-to-video result](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-reference-to-video/query-result.md): Check HappyHorse reference-to-video task status - [Wan 2.7 -- Generate from reference](https://docs.qwencloud.com/api-reference/video-generation/wan27-reference-to-video/create-task.md): Submit a Wan 2.7 reference-to-video task - [Wan 2.7 -- Retrieve reference-to-video result](https://docs.qwencloud.com/api-reference/video-generation/wan27-reference-to-video/query-result.md): Check Wan 2.7 video task status - [Wan 2.6 — Generate from reference](https://docs.qwencloud.com/api-reference/video-generation/wan-reference-to-video/create-task.md): Submit reference video - [Wan 2.6 — Retrieve reference-to-video result](https://docs.qwencloud.com/api-reference/video-generation/wan-reference-to-video/query-result.md): Check video task status - [HappyHorse -- Generate a video from text](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-text-to-video/create-task.md): Submit HappyHorse text-to-video task - [HappyHorse -- Retrieve text-to-video result](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-text-to-video/query-result.md): Check HappyHorse text-to-video task status - [Wan 2.7 -- Generate a video from text](https://docs.qwencloud.com/api-reference/video-generation/wan27-text-to-video/create-task.md): Submit text-to-video task (wan2.7) - [Wan 2.7 -- Retrieve text-to-video result](https://docs.qwencloud.com/api-reference/video-generation/wan27-text-to-video/query-result.md): Check wan2.7 video task status - [Wan 2.6 and earlier — Generate a video from text](https://docs.qwencloud.com/api-reference/video-generation/wan-text-to-video/create-task.md): Submit text-to-video task (wan2.6 and earlier) - [Wan 2.6 and earlier — Retrieve text-to-video result](https://docs.qwencloud.com/api-reference/video-generation/wan-text-to-video/query-result.md): Check video task status (wan2.6 and earlier) - [HappyHorse -- Edit a video](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-video-editing/create-task.md): Submit HappyHorse video editing task - [HappyHorse -- Retrieve video editing result](https://docs.qwencloud.com/api-reference/video-generation/happyhorse-video-editing/query-result.md): Check HappyHorse video editing task status - [Wan 2.7 -- Edit a video](https://docs.qwencloud.com/api-reference/video-generation/wan27-video-editing/create-task.md): Submit video editing task (wan2.7) - [Wan 2.7 -- Retrieve video editing result](https://docs.qwencloud.com/api-reference/video-generation/wan27-video-editing/query-result.md): Check wan2.7 video editing task status - [Wan — Edit a video](https://docs.qwencloud.com/api-reference/video-generation/wan-general-video-editing/create-task.md): Submit video edit task - [Wan — Retrieve video editing result](https://docs.qwencloud.com/api-reference/video-generation/wan-general-video-editing/query-result.md): Check video edit status - [Wan — Animate an image](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-animation/create-task.md): Submit animation task - [Wan — Retrieve animation result](https://docs.qwencloud.com/api-reference/video-generation/wan-image-to-animation/query-result.md): Check animation status - [Wan — Swap video character](https://docs.qwencloud.com/api-reference/video-generation/wan-video-character-swap/create-task.md): Submit face swap task - [Wan — Retrieve character swap result](https://docs.qwencloud.com/api-reference/video-generation/wan-video-character-swap/query-result.md): Check face swap status ## Speech-to-Text - [Fun-ASR WebSocket API](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/websocket-api.md): WebSocket connection, headers, and interaction flow for Fun-ASR real-time speech recognition - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition client events](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/client-events.md): WebSocket client event reference for Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime server events](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/server-events.md): WebSocket server event reference for Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime Python SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/python-sdk.md): This topic describes the parameters and interfaces of the Python SDK for the Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition model. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime Java SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/java-sdk.md): Real-time ASR Java SDK for Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime Android SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/android-sdk.md): This guide shows you how to use the Android SDK for Qwen-Audio-3.0-ASR-Flash-Streaming and Fun-ASR-Realtime real-time speech recognition to convert speech to text. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime iOS SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/ios-sdk.md): This guide shows you how to use the Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime real-time speech recognition iOS SDK to transcribe speech into text. - [Qwen-Audio-3.0-ASR-Flash-Streaming/Fun-ASR-Realtime HarmonyOS SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-realtime/harmonyos-sdk.md): Learn about HarmonyOS SDK integration, request parameters, APIs, callbacks, and sample code for real-time speech recognition. - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR recording Python SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/python-sdk.md): File transcription Python - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR HTTP API for non-real-time speech recognition](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/restful-api.md): File transcription REST - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR recording Java SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/java-sdk.md): File transcription Java SDK for Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR recording Android SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/android-sdk.md): File transcription Android SDK for Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR recording iOS SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/ios-sdk.md): File transcription iOS SDK for Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR - [Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR recording HarmonyOS SDK](https://docs.qwencloud.com/api-reference/speech-recognition/fun-asr-recording/harmonyos-sdk.md): File transcription HarmonyOS SDK for Qwen-Audio-3.0-ASR-Flash-Filetrans/Fun-ASR - [Qwen-ASR-Realtime WebSocket API](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr-realtime/websocket-api.md): WebSocket connection, headers, and interaction flows for Qwen-ASR-Realtime - [Qwen-ASR client events](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr-realtime/client-events.md): WebSocket client reference - [Qwen-ASR server events](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr-realtime/server-events.md): WebSocket server reference - [Qwen-ASR realtime Python SDK](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr-realtime/python-sdk.md): Qwen ASR Python streaming - [Qwen-ASR realtime Java SDK](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr-realtime/java-sdk.md): Qwen ASR Java integration - [OpenAI compatible ASR](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr/openai.md): Audio via Chat API - [DashScope synchronous](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr/dashscope.md): Sync audio recognition - [Qwen-ASR — Transcribe audio](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr/dashscope-async.md): Submit async transcription - [Qwen-ASR — Retrieve transcription](https://docs.qwencloud.com/api-reference/speech-recognition/qwen-asr/query-result.md): Check transcription status - [Custom hotwords HTTP API](https://docs.qwencloud.com/api-reference/speech-recognition/custom-hotwords/http-api.md): Manage custom vocabularies through HTTP APIs, including creating, listing, getting, updating, and deleting vocabularies. - [Custom hotwords Python SDK](https://docs.qwencloud.com/api-reference/speech-recognition/custom-hotwords/python-sdk.md): Use the Python SDK to create, list, query, update, and delete custom vocabularies for speech recognition. - [Custom hotwords Java SDK](https://docs.qwencloud.com/api-reference/speech-recognition/custom-hotwords/java-sdk.md): Use the Java SDK to create, query, update, and delete custom vocabularies for speech recognition. ## Text-to-speech - [Qwen-Audio-TTS/CosyVoice Java SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/java-sdk.md): Qwen-Audio-TTS/CosyVoice Java reference - [Qwen-Audio-TTS/CosyVoice Python SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/python-sdk.md): Qwen-Audio-TTS/CosyVoice Python reference - [CosyVoice WebSocket API](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/websocket-api.md): CosyVoice WebSocket API - [Qwen-Audio-TTS/CosyVoice client events](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/client-events.md): Qwen-Audio-TTS/CosyVoice real-time speech synthesis WebSocket client event reference - [CosyVoice server-side events](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/server-events.md): CosyVoice real-time speech synthesis WebSocket server event reference - [Qwen-Audio-TTS/CosyVoice Android SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/android-sdk.md): Convert text to expressive, high-quality speech in your Android app with the Qwen-Audio-TTS/CosyVoice text-to-speech (TTS) SDK. - [Qwen-Audio-TTS/CosyVoice iOS SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/ios-sdk.md): Use the Qwen-Audio-TTS/CosyVoice iOS SDK to turn text into high-quality, expressive speech in your iOS apps. - [Qwen-Audio-TTS/CosyVoice HarmonyOS SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/cosyvoice/harmonyos-sdk.md): Learn about HarmonyOS SDK integration, parameters, APIs, callbacks, and sample code for Qwen-Audio-TTS/CosyVoice real-time speech synthesis. - [Qwen-TTS client events](https://docs.qwencloud.com/api-reference/speech-synthesis/qwen-tts-realtime/client-events.md): Client events are JSON messages sent over a WebSocket connection to control the Qwen-TTS Realtime API session -- configure voice settings, stream text for synthesis, and signal completion. - [Qwen-TTS server events](https://docs.qwencloud.com/api-reference/speech-synthesis/qwen-tts-realtime/server-events.md): WebSocket server reference - [Qwen-TTS realtime Python SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/qwen-tts-realtime/python-sdk.md): Real-time TTS Python SDK - [Qwen-TTS realtime Java SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/qwen-tts-realtime/java-sdk.md): Real-time TTS Java SDK - [Speech synthesis](https://docs.qwencloud.com/api-reference/speech-synthesis/qwen-tts.md): Qwen-TTS API reference - [Create a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice/create-voice.md): Clone a voice from audio for Qwen-Audio-TTS or CosyVoice TTS models. Returns a `voice_id` that you can use in [Qwen-Audio-TTS](/developer-guides/speech/tts-models) or [CosyVoice synthesis calls](/api-reference/speech-synthesis/cosyvoice/python-sdk). - [List cloned voices](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice/list-voices.md): Returns a list of Qwen-Audio-TTS or CosyVoice cloned voices with optional prefix filtering and pagination. - [Query a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice/query-voice.md): Returns details about a specific Qwen-Audio-TTS or CosyVoice cloned voice, including its status, creation time, and the original audio URL. - [Update a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice/update-voice.md): Replace the audio of an existing Qwen-Audio-TTS or CosyVoice cloned voice. The voice ID remains the same. - [Delete a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice/delete-voice.md): Deletes a Qwen-Audio-TTS or CosyVoice cloned voice and releases its quota. - [Create a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/qwen/create-voice.md): Clone a voice from audio for Qwen TTS or Qwen-Omni-Realtime models. Returns a `voice` name that you can use in [Qwen TTS](/api-reference/speech-synthesis/qwen-tts), [Realtime streaming TTS](/api-reference/speech-synthesis/qwen-tts-realtime/client-events), or [Realtime Multimodal API](/api-reference/real-time-multimodal/client-events) calls. - [List cloned voices](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/qwen/list-voices.md): Returns a paginated list of Qwen cloned voices under your account. - [Delete a cloned voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/qwen/delete-voice.md): Deletes a Qwen cloned voice and releases its quota. - [Python SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice-sdk/python-sdk.md): Qwen-Audio-TTS/CosyVoice voice cloning Python SDK reference (VoiceEnrollmentService). - [Java SDK](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-cloning/cosyvoice-sdk/java-sdk.md): Qwen-Audio-TTS/CosyVoice voice cloning Java SDK reference (VoiceEnrollmentService). - [Create a designed voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/cosyvoice/create-voice.md): Design a voice from a text description for CosyVoice TTS models. Returns a `voice_id` and Base64-encoded preview audio. - [List designed voices](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/cosyvoice/list-voices.md): Returns a list of CosyVoice designed voices with optional prefix filtering and pagination. Uses the same `voice-enrollment` model as voice cloning — use the `prefix` filter to narrow results (designed voices follow the pattern `{target_model}-vd-{prefix}-{unique_id}`). - [Query a designed voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/cosyvoice/query-voice.md): Returns details about a specific CosyVoice designed voice, including its status, voice prompt, and preview text. - [Delete a designed voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/cosyvoice/delete-voice.md): Deletes a CosyVoice designed voice and releases its quota. - [Create a voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/qwen/create-voice.md): Create a custom voice from a text description and return preview audio. - [List designed voices](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/qwen/list-voices.md): Returns a paginated list of voices under your account. - [Query a voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/qwen/query-voice.md): Returns details about a specific voice. - [Delete a voice](https://docs.qwencloud.com/api-reference/speech-synthesis/voice-design/qwen/delete-voice.md): Deletes a voice and releases its quota. ## Speech-to-Speech - [Qwen-Omni client events](https://docs.qwencloud.com/api-reference/real-time-multimodal/client-events.md): WebSocket client reference - [Qwen-Omni server events](https://docs.qwencloud.com/api-reference/real-time-multimodal/server-events.md): WebSocket server reference - [Qwen-Omni Python SDK](https://docs.qwencloud.com/api-reference/real-time-multimodal/realtime-python-sdk.md): Qwen-Omni-Realtime Python SDK - [Qwen-Omni Java SDK](https://docs.qwencloud.com/api-reference/real-time-multimodal/realtime-java-sdk.md): Qwen-Omni-Realtime Java SDK - [Qwen-Audio-Realtime WebSocket API](https://docs.qwencloud.com/api-reference/qwen-audio-realtime/websocket-api.md): Qwen-Audio-Realtime WebSocket connection protocol, request headers, core concepts, and interaction flows - [Qwen-Audio-Realtime Client Events](https://docs.qwencloud.com/api-reference/qwen-audio-realtime/client-events.md): Qwen-Audio-Realtime API client event reference - [Qwen-Audio-Realtime Server Events](https://docs.qwencloud.com/api-reference/qwen-audio-realtime/server-events.md): Qwen-Audio-Realtime API server event reference - [LiveTranslate client events](https://docs.qwencloud.com/api-reference/speech-translation/livetranslate-realtime/client-events.md): WebSocket client reference - [LiveTranslate server events](https://docs.qwencloud.com/api-reference/speech-translation/livetranslate-realtime/server-events.md): WebSocket server reference - [LiveTranslate Python SDK](https://docs.qwencloud.com/api-reference/speech-translation/livetranslate-realtime/python-sdk.md): LiveTranslate Python SDK - [LiveTranslate Java SDK](https://docs.qwencloud.com/api-reference/speech-translation/livetranslate-realtime/java-sdk.md): LiveTranslate Java SDK - [Audio and video translation](https://docs.qwencloud.com/api-reference/speech-translation/audio-video-translation-api.md): LiveTranslate API reference ## Realtime API - [Realtime API overview](https://docs.qwencloud.com/api-reference/realtime-api/overview.md): The Realtime API provides multiple transport protocols, each optimized for different requirements such as performance, latency, weak-network resilience, and integration cost - [SDK download](https://docs.qwencloud.com/api-reference/realtime-api/sdk-download.md): Download AOQ SDK and WebSocket SDK - [Token authentication](https://docs.qwencloud.com/api-reference/realtime-api/token-auth.md): Learn how the Realtime API authenticates connections with tokens, including how to get an API key and how to authenticate over the WebSocket, WebRTC, and AOQ protocols - [Connect to a model](https://docs.qwencloud.com/api-reference/realtime-api/connect-model.md): Connect to Realtime API models and applications over the AOQ, WebRTC, and WebSocket protocols. Covers the connection flow, sequence diagrams, and code examples for each protocol - [SDK overview](https://docs.qwencloud.com/api-reference/realtime-api/aoq-sdk-intro.md): AOQ SDK platform support and architecture overview - [Android SDK](https://docs.qwencloud.com/api-reference/realtime-api/android-sdk.md): AOQ Android SDK API reference - [iOS SDK](https://docs.qwencloud.com/api-reference/realtime-api/ios-sdk.md): AOQ iOS SDK API reference - [HarmonyOS SDK](https://docs.qwencloud.com/api-reference/realtime-api/harmonyos-sdk.md): AOQ HarmonyOS SDK API reference - [Windows SDK](https://docs.qwencloud.com/api-reference/realtime-api/windows-sdk.md): AOQ Windows SDK C++ API reference - [macOS SDK](https://docs.qwencloud.com/api-reference/realtime-api/macos-sdk.md): AOQ macOS SDK Objective-C API reference - [Electron SDK](https://docs.qwencloud.com/api-reference/realtime-api/electron-sdk.md): AOQ Client SDK Electron API reference - [Linux Python SDK](https://docs.qwencloud.com/api-reference/realtime-api/linux-python-sdk.md): AOQ Client SDK Linux Python API reference - [Linux C++ SDK](https://docs.qwencloud.com/api-reference/realtime-api/linux-cpp-sdk.md): AOQ Client SDK Linux C++ API reference - [Connection management](https://docs.qwencloud.com/api-reference/realtime-api/connection-management.md): AOQ SDK connection state diagram, migration rules, and API - [Media stream control](https://docs.qwencloud.com/api-reference/realtime-api/media-stream-control.md): AOQ SDK media stream send/stop control - [Audio features](https://docs.qwencloud.com/api-reference/realtime-api/audio-features.md): AOQ SDK audio capture, playback, and codec configuration - [Custom audio playback](https://docs.qwencloud.com/api-reference/realtime-api/custom-audio-playback.md): Play audio returned by AOQ using an external player - [Custom audio capture](https://docs.qwencloud.com/api-reference/realtime-api/custom-audio-capture.md): Use an external audio source instead of the device microphone - [Video features](https://docs.qwencloud.com/api-reference/realtime-api/video-features.md): AOQ SDK video capture, preview, and encoding configuration - [Custom video input](https://docs.qwencloud.com/api-reference/realtime-api/custom-video-input.md): Use an external video source instead of the device camera ## Text Embeddings - [OpenAI compatible embedding](https://docs.qwencloud.com/api-reference/text-embedding/openai-embedding.md): OpenAI-compatible embedding - [DashScope embedding](https://docs.qwencloud.com/api-reference/text-embedding/dashscope-embedding.md): DashScope embedding API ## Multimodal Embeddings - [DashScope multimodal embedding](https://docs.qwencloud.com/api-reference/multimodal-embedding/dashscope-multimodal-embedding.md): Multimodal embedding API ## Reranking - [OpenAI compatible reranking](https://docs.qwencloud.com/api-reference/rerank/openai-rerank.md): OpenAI-compatible reranking API - [DashScope reranking](https://docs.qwencloud.com/api-reference/rerank/dashscope-rerank.md): DashScope reranking API ## Platform API - [Conversations](https://docs.qwencloud.com/api-reference/platform-api/conversations.md): Auto-managed chat history - [Create conversation](https://docs.qwencloud.com/api-reference/platform-api/conversations/create-conversation.md): Create a conversation with optional initial message items. - [Retrieve conversation](https://docs.qwencloud.com/api-reference/platform-api/conversations/retrieve-conversation.md): Retrieve a conversation by its ID. - [Update conversation](https://docs.qwencloud.com/api-reference/platform-api/conversations/update-conversation.md): Update a conversation's metadata. This completely replaces existing metadata. - [Delete conversation](https://docs.qwencloud.com/api-reference/platform-api/conversations/delete-conversation.md): Delete a conversation. Message items are not deleted. - [Create items](https://docs.qwencloud.com/api-reference/platform-api/conversations/create-items.md): Add message items to a conversation. - [List items](https://docs.qwencloud.com/api-reference/platform-api/conversations/list-items.md): List message items in a conversation. - [Retrieve item](https://docs.qwencloud.com/api-reference/platform-api/conversations/retrieve-item.md): Retrieve a message item by its ID. - [Delete item](https://docs.qwencloud.com/api-reference/platform-api/conversations/delete-item.md): Delete a message item from a conversation. - [File](https://docs.qwencloud.com/api-reference/platform-api/file.md): File management API - [Upload file](https://docs.qwencloud.com/api-reference/platform-api/file/upload-file.md): Upload a file for document analysis or batch processing. You can store up to 10,000 files and 100 GB total. Files never expire. - [Retrieve file](https://docs.qwencloud.com/api-reference/platform-api/file/retrieve-file.md): Retrieve file details by file ID. - [List files](https://docs.qwencloud.com/api-reference/platform-api/file/list-files.md): List all files in your account (uploaded files and batch results), with filtering by purpose, creation time, and pagination. - [Delete file](https://docs.qwencloud.com/api-reference/platform-api/file/delete-file.md): Delete a file by its ID. - [Create batch](https://docs.qwencloud.com/api-reference/platform-api/batch/create-batch.md): Create a batch job that processes all requests in the uploaded input file asynchronously. - [Retrieve batch](https://docs.qwencloud.com/api-reference/platform-api/batch/retrieve-batch.md): Check the status of a batch job and retrieve output file IDs when the job completes. - [List batches](https://docs.qwencloud.com/api-reference/platform-api/batch/list-batches.md): List batch jobs for your account. Results are sorted by creation time in descending order. Only tasks from the last 30 days are available. - [Cancel batch](https://docs.qwencloud.com/api-reference/platform-api/batch/cancel-batch.md): Cancel an in-progress or queued batch job. The status changes to `cancelling` while currently executing requests complete, then to `cancelled`. Completed requests before cancellation are still billed. ## Toolkit & Framework - [OpenAI compatibility](https://docs.qwencloud.com/api-reference/toolkitframework/openai-compatible/overview.md): Migrate from OpenAI by changing three parameters: base_url, api_key, and model. ## More - [Generate a temporary API key](https://docs.qwencloud.com/api-reference/more/generate-a-temporary-api-key.md): Short-lived access tokens - [Manage asynchronous tasks](https://docs.qwencloud.com/api-reference/more/manage-asynchronous-tasks.md): Track and control async jobs ## Billing - [Billing overview](https://docs.qwencloud.com/resources/billing-overview.md): View bills, manage payments, and track spending - [Free quota](https://docs.qwencloud.com/resources/free-quota.md): New user free quota - [Use and manage coupons](https://docs.qwencloud.com/resources/coupons.md): View, use, and manage your coupons - [Billing and cost management](https://docs.qwencloud.com/resources/bill-query.md): Analyze and manage costs - [Overdue payment protection](https://docs.qwencloud.com/resources/overdue-payment-protection.md): Understand grace periods for overdue payments ## FAQ - [Account & access](https://docs.qwencloud.com/resources/faq-account.md): Keys, permissions, access - [Billing & pricing](https://docs.qwencloud.com/resources/faq-billing.md): Payments and costs Q&A - [Text generation FAQ](https://docs.qwencloud.com/resources/faq-text-generation.md): Common questions about streaming, context management, context caching, structured output, thinking mode, function calling, and batch. - [Images & videos FAQ](https://docs.qwencloud.com/resources/faq-images-videos.md): Common questions about image and video generation — billing, API errors, model differences, input requirements, and output URLs. - [Audio & speech FAQ](https://docs.qwencloud.com/resources/faq-audio-speech.md): CosyVoice TTS, Qwen-Omni Realtime, and Fun-ASR — common questions about synthesis, real-time conversation, and speech recognition. - [Embedding & reranking FAQ](https://docs.qwencloud.com/resources/faq-embedding-reranking.md): Common questions about text embeddings, multimodal embeddings, and reranking — model selection, dimensions, batch limits, and use cases. ## Changelog - [Model releases](https://docs.qwencloud.com/changelog/models.md): New launches and updates - [Platform updates](https://docs.qwencloud.com/changelog/platform.md): Features and improvements - [Legal updates](https://docs.qwencloud.com/changelog/legal-updates.md) ## OpenAPI Specs - [openapi-anthropic](https://docs.qwencloud.com/openapi-anthropic.json): Call Qwen models through the Anthropic-compatible Messages API. Supports thinking, tool use, streaming, image/video understanding, and context caching. Authentication: pass your API key via the `x-api-key` header or the `Authorization: Bearer` header. - [openapi-dashscope](https://docs.qwencloud.com/openapi-dashscope.json): Call Qwen models using the DashScope native HTTP API. Supports text and multimodal models, streaming, tool calls, and structured output. - [openapi-openai-chat](https://docs.qwencloud.com/openapi-openai-chat.json): Call Qwen models using the OpenAI-compatible Chat API. Supports text and multimodal models, streaming, tool calls, and structured output. - [openapi-openai-responses](https://docs.qwencloud.com/openapi-openai-responses.json): Call Qwen models using the OpenAI-compatible Responses API. Supports built-in tools, multi-turn context management, and thinking mode. - [openapi-qwen-image-edit](https://docs.qwencloud.com/openapi-qwen-image-edit.json): Qwen-Image editing API. Supports single-image editing, multi-image fusion, style transfer, text editing, element manipulation, and more. - [openapi-qwen-image](https://docs.qwencloud.com/openapi-qwen-image.json): Qwen-Image text-to-image generation API. Supports synchronous calls for all models and asynchronous calls for legacy models. - [openapi-wan-t2i-v2](https://docs.qwencloud.com/openapi-wan-t2i-v2.json): Generate images from text prompts using the Wan text-to-image model family. Supports various artistic styles and realistic photographic effects to meet diverse creative needs. This API uses an asynchronous task pattern: submit a task with POST, then poll for results with GET. - [openapi-wan25-image-edit](https://docs.qwencloud.com/openapi-wan25-image-edit.json): Wan2.5 general image editing API. Edits images using text prompts, preserving subject consistency. Supports single-image editing and multi-image fusion with up to three reference images. - [openapi-wan26-image](https://docs.qwencloud.com/openapi-wan26-image.json): Wan2.6 image generation and editing API. Supports multi-image input, image editing, and interleaved text-image output. - [openapi-wan27-image](https://docs.qwencloud.com/openapi-wan27-image.json): Wan2.7 image generation and editing API. Supports text-to-image, multi-image editing, interactive editing with bounding boxes, and image set generation. - [openapi-z-image](https://docs.qwencloud.com/openapi-z-image.json): Z-Image text-to-image generation API. - [openapi-image-translation](https://docs.qwencloud.com/openapi-image-translation.json): Qwen-MT-Image accurately translates text in images while preserving the original layout. It supports domain hints, sensitive word filtering, and terminology intervention. `qwen-mt-image-2.0` supports both synchronous and asynchronous modes. See [Synchronous calls](/api-reference/image-translation/qwen-mt-image/synchronous) for the sync API. Asynchronous workflow: 1. **Create task**: POST to submit a task and receive a task_id. 2. **Query result**: Poll by task_id until the task completes and returns the image URL. - [openapi-image](https://docs.qwencloud.com/openapi-image.json): DashScope Image APIs for image generation, editing, specialized image tasks, and multimodal embeddings. - [openapi-batch](https://docs.qwencloud.com/openapi-batch.json): Submit batch inference jobs via file upload at 50% of real-time API cost. OpenAI-compatible. - [openapi-conversations](https://docs.qwencloud.com/openapi-conversations.json): Auto-managed chat history for multi-turn context across devices and sessions. - [openapi-file](https://docs.qwencloud.com/openapi-file.json): Upload and manage files for document analysis or batch processing. - [openapi-reranking](https://docs.qwencloud.com/openapi-reranking.json): Rerank retrieved documents by semantic relevance to improve search accuracy in RAG and retrieval systems. - [openapi-qwen-asr](https://docs.qwencloud.com/openapi-qwen-asr.json): API reference for Qwen-ASR audio file recognition. Supports OpenAI-compatible, DashScope synchronous, and DashScope asynchronous protocols. - [openapi-qwen-tts](https://docs.qwencloud.com/openapi-qwen-tts.json): API reference for the Qwen speech synthesis model (Qwen-TTS, Qwen3-TTS-Flash, Qwen3-TTS-Instruct-Flash). Supports streaming and non-streaming output. - [openapi-voice-cloning-cosyvoice](https://docs.qwencloud.com/openapi-voice-cloning-cosyvoice.json) - [openapi-voice-cloning-cosyvoice-delete](https://docs.qwencloud.com/openapi-voice-cloning-cosyvoice-delete.json) - [openapi-voice-cloning-cosyvoice-list](https://docs.qwencloud.com/openapi-voice-cloning-cosyvoice-list.json) - [openapi-voice-cloning-cosyvoice-query](https://docs.qwencloud.com/openapi-voice-cloning-cosyvoice-query.json) - [openapi-voice-cloning-cosyvoice-update](https://docs.qwencloud.com/openapi-voice-cloning-cosyvoice-update.json) - [openapi-voice-cloning-qwen](https://docs.qwencloud.com/openapi-voice-cloning-qwen.json) - [openapi-voice-cloning-qwen-delete](https://docs.qwencloud.com/openapi-voice-cloning-qwen-delete.json) - [openapi-voice-cloning-qwen-list](https://docs.qwencloud.com/openapi-voice-cloning-qwen-list.json) - [openapi-voice-design-cosyvoice](https://docs.qwencloud.com/openapi-voice-design-cosyvoice.json) - [openapi-voice-design-cosyvoice-delete](https://docs.qwencloud.com/openapi-voice-design-cosyvoice-delete.json) - [openapi-voice-design-cosyvoice-list](https://docs.qwencloud.com/openapi-voice-design-cosyvoice-list.json) - [openapi-voice-design-cosyvoice-query](https://docs.qwencloud.com/openapi-voice-design-cosyvoice-query.json) - [openapi-voice-design](https://docs.qwencloud.com/openapi-voice-design.json): Create and manage custom voices from text descriptions for use with Qwen TTS models. - [openapi-voice-design-delete](https://docs.qwencloud.com/openapi-voice-design-delete.json) - [openapi-voice-design-list](https://docs.qwencloud.com/openapi-voice-design-list.json) - [openapi-voice-design-query](https://docs.qwencloud.com/openapi-voice-design-query.json) - [openapi-text-embedding](https://docs.qwencloud.com/openapi-text-embedding.json): API reference for text embedding models. Supports OpenAI-compatible and DashScope protocols. - [openapi-happyhorse-image-to-video](https://docs.qwencloud.com/openapi-happyhorse-image-to-video.json): Generate videos from an image using the HappyHorse model. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-happyhorse-ref-to-video](https://docs.qwencloud.com/openapi-happyhorse-ref-to-video.json): The HappyHorse reference-to-video model lets you provide multiple reference images and a text prompt to generate a video that combines subjects from the images into a scene based on the prompt. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-happyhorse-text-to-video](https://docs.qwencloud.com/openapi-happyhorse-text-to-video.json): Generate videos from text using the HappyHorse model. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-happyhorse-video-editing](https://docs.qwencloud.com/openapi-happyhorse-video-editing.json): Edit videos using the HappyHorse model with text instructions and optional reference images. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-wan-video-editing](https://docs.qwencloud.com/openapi-wan-video-editing.json): Wan unified video editing API. Supports multimodal input (text, image, video) with five core capabilities: multi-image reference, video repainting, local editing, video extension, and frame expansion. - [openapi-wan-image-to-animation](https://docs.qwencloud.com/openapi-wan-image-to-animation.json): Animate a character image by transferring actions and expressions from a reference video using the wan2.2-animate-move model. - [openapi-wan-i2v-first-frame](https://docs.qwencloud.com/openapi-wan-i2v-first-frame.json): Generate videos from a first-frame image and text prompt using Wan image-to-video models. Supports audio synchronization, multi-shot narrative, and multiple resolution tiers. - [openapi-wan-i2v-first-last](https://docs.qwencloud.com/openapi-wan-i2v-first-last.json): Generate a smoothly transitioning video from a first frame image, a last frame image, and a text prompt using the Wan kf2v model. - [openapi-wan-ref-to-video](https://docs.qwencloud.com/openapi-wan-ref-to-video.json): Wan reference-to-video API. Generates performance videos from reference images or videos with multimodal input (text, image, video). Supports single or multi-character interactions, multi-shot narrative, and audio-video sync. - [openapi-wan-text-to-video](https://docs.qwencloud.com/openapi-wan-text-to-video.json): Wan text-to-video generation API. Supports multimodal input (text, images, audio) and generates videos up to 15 seconds at 1080P resolution. Uses asynchronous task submission — submit a task, then poll for results. - [openapi-wan-character-swap](https://docs.qwencloud.com/openapi-wan-character-swap.json): Wan video character swap API. Replaces the main character in a video with a character from an image, preserving the scene, lighting, and tone of the original video. Uses asynchronous task processing. - [openapi-wan27-image-to-video](https://docs.qwencloud.com/openapi-wan27-image-to-video.json): The Wan 2.7 image-to-video model now supports multi-modal input — including images, audio, and video — and performs three main tasks: video generation from the first frame, video generation from the first and last frames, and video continuation. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-wan27-ref-to-video](https://docs.qwencloud.com/openapi-wan27-ref-to-video.json): Wan 2.7 reference-to-video API. Generates performance videos from reference images or videos with multimodal input (text, image, video). Uses the new protocol with media array input, resolution+ratio parameters, and enhanced response metadata. - [openapi-wan27-text-to-video](https://docs.qwencloud.com/openapi-wan27-text-to-video.json): Generate videos from text using the Wan 2.7 model. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-wan27-video-editing](https://docs.qwencloud.com/openapi-wan27-video-editing.json): Edit videos using the Wan 2.7 model. Supports multimodal inputs such as text, images, and videos for video editing. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result. - [openapi-wan30-video](https://docs.qwencloud.com/openapi-wan30-video.json): Wan 3.0 is an All-in-One video generation model that supports text-to-video, image-to-video (first-frame / first+last-frame), reference-to-video, and file-to-video, generating videos up to 30 seconds. Submit a task asynchronously, then poll `GET /tasks/{task_id}` for the result.