New launches and updates
Some legacy models are being retired soon. See Model deprecation policy for the deprecation timeline and replacement models. For model upgrade, version switch, update, and pricing change announcements, see Model update announcements for details.
September 2, 2026
qwen3.8-max-0902
Qwen3.8-Max-0902(alias qwen3.8-max-2026-09-02)is an upgraded snapshot of qwen3.8-max. Coding capability breaks new ground, handling more complex engineering-scale projects and long-horizon autonomous development. Collaborative agent performance is significantly enhanced, with greater composure in multi-tool orchestration and end-to-end task delivery. Native vision understanding is refined across chart reasoning, document parsing, and multimodal perception — sharper and more reliable. Retains the 1M context window, thinking mode, and full tool ecosystem, evolving at a higher level of intelligence. See Text generation models.August 30, 2026
qwen-flash-character
The Qwen Role-Playing Model Series is specifically optimized for multi-language anthropomorphic interaction scenarios. It demonstrates advanced capabilities in character consistency maintenance, context-aware dialogue progression, and empathetic engagement, enabling precise personalized character embodiment. Developer guideAugust 28, 2026
qwen-mt-image-2.0
A model service dedicated to image translation, capable of translating images across 55 languages including Chinese, English, and Japanese. It accurately preserves image layout and content information, and supports terminology customization, sensitive word filtering, and product subject detection for flexible, accurate, and efficient image localization. API referenceAugust 26, 2026
qwen3.8-flash
Qwen3.8-Flash is the latest multimodal model from the Qwen family, combining powerful reasoning and generation with remarkable speed. It natively supports a million-token context window, allowing it to process lengthy documents, entire codebases, and complex conversations in a single pass. It shines in coding assistance, agentic workflows, and visual understanding — whether it's fixing code autonomously, operating desktop applications, or analyzing charts and long videos. Fully compatible with both OpenAI and Anthropic API protocols, it integrates seamlessly with popular developer tools like Claude Code and Codex, making it easy to build high-concurrency applications and intelligent workflows. With strong performance and highly competitive inference costs, Qwen3.8-Flash is an ideal choice for developers and businesses seeking the best of both worlds in AI applications. See Text generation models.August 24, 2026
wan3.0-video General Availability
Wan 3.0 is an All-in-One video generation model that supports text-to-video, image-to-video (first-frame / first+last-frame), and reference-to-video in a single unified model. Features include audio generation, sound toggle, smart duration, and adaptive aspect ratio. Initially launched as invite-only on August 6, now publicly available. API referencewan3.0-video-prime
Wan3.0-Video-Prime is the high-speed variant of Wan 3.0, offering the same All-in-One capabilities as wan3.0-video with faster generation speed. API referenceAugust 19, 2026
qwen3.8-27b
The Qwen3.8 27B native vision-language dense model builds upon the 3.6-27B version, with key improvements in coding and office productivity capabilities across both text and visual modalities. It enables more reliable end-to-end completion of complex tasks, delivering consistently trustworthy results. See Text generation models.ZHIPU/GLM-5.3
The Z.AI direct-supply GLM-5.3 model is now available. GLM-5.3 is Zhipu AI's most powerful model for programming capabilities to date, demonstrating a 50% improvement over GLM-5.2 in internal subjective evaluations. In terms of cybersecurity, GLM-5.3 performs on par with Mythos 5 in tasks such as white-box code review and vulnerability discovery, showcasing its strong potential for cybersecurity defense scenarios. User guidekimi-k3
Kimi K3 model is now available in the Singapore region. Kimi's most capable flagship model to date, it always reasons and uses preserved thinking. Supports text and image input, conversation and agent tasks, and dynamic tool loading. User guideAugust 13, 2026
deepseek-v4-pro-0813
A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis. Supports streaming, thinking mode, Function Calling, structured output, web search, and context caching. User guideqwen3.8-2.4t-a95b
Qwen3.8-2.4T-A95B is the open-source version of the latest Qwen flagship series, featuring a sparse MoE architecture with 2.4 trillion total parameters and approximately 95 billion activated per step. Combined with hybrid attention mechanisms, it supports a 1 million token context window. Thinking mode is enabled by default. Significant improvements in coding, productivity, research, and long-horizon agent tasks. See Text generation models.August 11, 2026
qwen3.7-text-embedding
Qwen3.7-Text-Embedding is a multilingual text vector model based on Qwen3.7, offering significant improvements in text retrieval, clustering, and classification over text-embedding-v4. It achieves 20% better performance in MTEB multilingual, Chinese-English, and code retrieval tasks and supports customizable vector dimensions (256-2560). User guideAugust 10, 2026
qwen-audio-3.0-realtime-plus, qwen-audio-3.0-realtime-flash
Qwen-Audio end-to-end real-time speech model balances speech reasoning and duplex conversation rhythm. It maintains fluid, natural real-time interactions while effectively controlling end-to-end response latency through parallel inference and omni-directional streaming optimizations. User guideZHIPU/GLM-5.2
The Z.AI direct-supply GLM-5.2 model is now available. User guideAugust 6, 2026
wan3.0-video
Wan 3.0 All-in-One video generation model, initially launched for invite-only testing (now publicly available). Supports text-to-video, image-to-video (first-frame / first+last-frame), and reference-to-video in a single unified model. API referenceAugust 4, 2026
qwen-image-3.0-pro, qwen-image-3.0
The Qwen-Image-3.0-Pro model supports long-text input and dense image-in-image layouts, enabling precise one-shot generation of complex layouts such as newspapers, storyboards, menus, and exam papers. It features 10-pixel fine-grained text rendering, photographic-level detail reproduction of micro-expressions, pores and hair strands, and supports 12 languages, multiple fonts, and high-fidelity rendering of mainstream web and game interfaces. qwen-image-3.0 is the standard variant, balancing quality and speed.August 3, 2026
qwen3.8-max
The Qwen3.8 native vision-language Max series model is Qwen's most capable flagship model to date, featuring 2.4 trillion parameters with a Mixture-of-Experts architecture. It delivers outstanding performance comparable to the current state-of-the-art models, with significant improvements over the 3.7 series. Supports hybrid thinking mode (enabled by default) and a 1M-token context window. See Text generation models.August 1, 2026
deepseek-v4-flash-0731
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. User guideJuly 30, 2026
qwen-audio-3.0-asr-flash-streaming, qwen-audio-3.0-asr-flash-filetrans, qwen-audio-3.0-asr-flash
Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) speech recognition models:- Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents.
- Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios.
- Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats.
- Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean.
- Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms.
July 25, 2026
qwen3.7-flash, qwen3.7-flash-2026-07-15
The Qwen3.7 native vision-language Flash model series delivers a comprehensive upgrade over 3.6-Flash in multimodal understanding and agent execution. This model particularly excels in enhanced multimodal foundations with stronger universal object recognition, further improved real-world perception and spatial intelligence, significantly upgraded multimodal agent capabilities for Search Agent and CI Agent scenarios with more stable end-to-end task execution, as well as optimized multimodal coding for a smoother vibe coding experience.July 14, 2026
qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash
The Qwen-Audio-TTS speech synthesis model is now available. It adds support for more minority languages and Chinese dialects, with enhanced instruction following and fine-grained tag control. Audio quality and expressiveness are significantly improved. The Plus edition targets high-quality professional scenarios, while the Flash edition targets low-latency real-time interaction. User guideJuly 1, 2026
wan2.7-t2v-2026-06-12
A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same as wan2.7-t2v. API referencewan2.7-r2v-2026-06-12
A snapshot version of the Wan 2.7 reference-to-video model. It supports subject referencing and voice customization, and can generate a scripted video directly from a single, multi-panel storyboard image. API referenceJune 26, 2026
glm-5.2
The Zhipu GLM-5.2 model is the latest flagship designed for long-horizon tasks, supporting an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance.kimi-k2.7-code
Kimi K2.7 Code model is now available in the Singapore region. An agent-centric coding model optimized for long-range software engineering tasks. Supports thinking mode only.June 25, 2026
qwen-image-2.0-pro-2026-06-22
The latest snapshot of the Qwen-Image-2.0 series, unifying image generation and editing. Compared with the April 22 snapshot, this model features enhanced text rendering with support for up to 1k token instructions, more refined photorealistic quality and scene detail, and stronger semantic adherence.June 22, 2026
HappyHorse 1.1
happyhorse-1.1-t2v
HappyHorse 1.1 text-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.1-i2v
HappyHorse 1.1 image-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.1-r2v
HappyHorse 1.1 reference-to-video model. Supports multiple reference images as input to generate audio videos with 3--15 seconds duration at 720P/1080P resolution. API referenceJune 17, 2026
fun-asr-flash-2026-06-15
Fun-ASR snapshot (June 2026 update). Adds full support for the seven major Chinese dialect systems (Mandarin, Wu, Xiang, Gan, Hakka, Min, Yue) and 20+ regional Mandarin accents, with specific optimizations for classical Chinese poetry recognition. Improves punctuation prediction and text normalization for numbers, dates, and amounts. Language support expanded to 30 languages. Supports contextualization and can transcribe audio up to 5 minutes.June 10, 2026
qwen3.7-max-2026-06-08
The Max model, the largest and most capable in the Qwen3.7 series, has added visual-modal understanding compared to the May 20 snapshot, enabling it to perceive real-world scenes and supporting multimodal interactive hybrid agent capabilities.June 1, 2026
qwen3.7-plus, qwen3.7-plus-2026-05-26
Qwen 3.7 Plus series builds upon strong text capabilities with comprehensively upgraded vision-language abilities while maintaining full agentic capabilities in coding, tool use, and productivity workflows. Its core differentiator is multimodal interactive hybrid agent capabilities — perceiving real-world scenarios, reading screens and operating GUIs, generating code based on visual references, and end-to-end navigation of mobile applications.May 27, 2026
glm-5.1
The Zhipu GLM-5.1 model is designed for long-horizon tasks and supports a 200K context window with a maximum output of 128K tokens. With strong logical reasoning, long-text comprehension, and code generation capabilities, it delivers excellent results across multiple benchmarks and is well suited for intelligent interaction, enterprise applications, and development assistance.May 25, 2026
qwen3.7-max-preview, qwen3.7-max-2026-05-17
Qwen Max series model snapshots. Text-only input, thinking mode enabled by default.May 22, 2026
qwen3.5-livetranslate-flash-realtime, qwen3.5-livetranslate-flash-realtime-2026-05-19
High-precision, real-time multilingual audio and video translation model built on Qwen3.5-Omni. Supports 60 languages (29 with audio + text output, 31 with text-only output), voice cloning for speaker-preserving translation, and visual context from video streams to improve accuracy. Upgrades from qwen3-livetranslate-flash-realtime with broader language coverage and lower latency. User guide | API referenceMay 21, 2026
qwen3.7-max, qwen3.7-max-2026-05-20
Next-generation flagship in the Qwen Max series. Text-only input, thinking mode enabled by default, supports explicit cache. Excels at coding, office & productivity, and long-horizon autonomous execution.May 11, 2026
deepseek-v4-pro, deepseek-v4-flash
DeepSeek V4 series. deepseek-v4-pro is a large-scale MoE model with strong general reasoning. deepseek-v4-flash is a lightweight, cost-effective model optimized for speed. Both support function calling, context cache, and thinking mode (enabled by default) with 1M context and a shared 384k output budget.April 27, 2026
HappyHorse video series
happyhorse-1.0-t2v
HappyHorse text-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-i2v
HappyHorse image-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-r2v
HappyHorse reference-to-video model. Supports multiple reference images as input to generate audio videos with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-video-edit
HappyHorse video editing model. Supports video editing and processing. API referenceApril 26, 2026
wan2.7-t2v-2026-04-25
A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same aswan2.7-t2v. API referencewan2.7-i2v-2026-04-25
A snapshot version of the Wan 2.7 image-to-video model. The model capabilities are the same aswan2.7-i2v. API referenceApril 23, 2026
qwen3.5-plus-2026-04-20
New Qwen3.5-Plus snapshot with significantly improved agentic coding and faster inference compared to the February 15 snapshot. Knowledge, reasoning, and long-context capabilities remain strong — ideal for coding agents, production workflows, and high-throughput scenarios.qwen3.6-27b
Qwen3.6 27B dense open-source vision-language model. Enhanced agentic coding and STEM reasoning over Qwen3.5-27B. Significant improvements in spatial intelligence, object localization and detection. Steady gains in video understanding, document OCR, and vision agent capabilities. Supports function calling and structured output but not built-in tools.qwen-image-2.0-pro-2026-04-22
New Qwen-Image-2.0-Pro snapshot with unified image generation and editing. Compared to the March 3 snapshot, this model delivers noticeable improvements in visual quality — especially texture detail, lighting, and materials. Supports multilingual in-image text rendering and more balanced artistic style expression.April 20, 2026
qwen3.6-max-preview
The largest closed-source model in the Qwen3.6 series with improved coding and more efficient Agent execution. Supports text-only input, thinking mode (enabled by default), explicit cache, and Function Calling.April 16, 2026
qwen3.6-flash, qwen3.6-flash-2026-04-16, qwen3.6-35b-a3b
Qwen3.6 native vision-language Flash series with significant overall improvements over Qwen3.5-Flash. Enhanced agent programming capabilities with major benchmark gains over the previous generation, stronger math and code reasoning, and improved spatial intelligence — particularly object localization and detection.fun-asr, fun-asr-2025-11-07
Fun-ASR real-time speech recognition upgrades: dialect support covering seven major Chinese dialect groups and 20+ regional accents, improved classical poetry recognition, enhanced punctuation prediction and text normalization (numbers, dates, monetary values), and multilingual support for 30 languages.April 3, 2026
Wan 2.7 video series
wan2.7-videoedit
Instruction-based video editing and migration. Supports content replacement with reference images and replicating actions, effects, and camera movements.wan2.7-i2v
Multimodal input (text, image, audio, video) for first-frame, start-and-end-frame, and video continuation tasks.wan2.7-t2v
New resolution options and custom aspect ratios for different scenarios and platforms.wan2.7-r2v
Entity reference, voice customization, and playbook-based video generation from a single storyboard.April 2, 2026
qwen3.6-plus, qwen3.6-plus-2026-04-02
Upgraded coding (Agentic Coding, frontend, Vibe Coding), multimodal (object recognition, OCR, localization), and general reasoning. Fixes known issues from Qwen-3.5-Plus.April 1, 2026
wan2.7-image-pro, wan2.7-image
Text-to-image, image-to-image, image editing, multi-image reference, and interactive editing. Pro series supports 4K output.March 31, 2026