New launches and updates
Some legacy models are being retired soon. See Model deprecation policy for the deprecation timeline and replacement models.
August 13, 2026
deepseek-v4-pro-0813
A flagship Mixture-of-Experts (MoE) large language model with 1.6 trillion total parameters and 49 billion activated parameters. It natively supports context windows of up to 1 million tokens. Trained on extensive high-quality data, the model delivers strong performance in mathematical and logical reasoning, complex reasoning, professional code generation, and in-depth long-document analysis. Supports streaming, thinking mode, Function Calling, structured output, web search, and context caching. User guideqwen3.8-2.4t-a95b
Qwen3.8-2.4T-A95B is the open-source version of the latest Qwen flagship series, featuring a sparse MoE architecture with 2.4 trillion total parameters and approximately 95 billion activated per step. Combined with hybrid attention mechanisms, it supports a 1 million token context window. Thinking mode is enabled by default. Significant improvements in coding, productivity, research, and long-horizon agent tasks. See Text generation models.August 11, 2026
qwen3.7-text-embedding
Qwen3.7-Text-Embedding is a multilingual text vector model based on Qwen3.7, offering significant improvements in text retrieval, clustering, and classification over text-embedding-v4. It achieves 20% better performance in MTEB multilingual, Chinese-English, and code retrieval tasks and supports customizable vector dimensions (256-2560). User guideAugust 10, 2026
qwen-audio-3.0-realtime-plus, qwen-audio-3.0-realtime-flash
Qwen-Audio end-to-end real-time speech model balances speech reasoning and duplex conversation rhythm. It maintains fluid, natural real-time interactions while effectively controlling end-to-end response latency through parallel inference and omni-directional streaming optimizations. User guideZHIPU/GLM-5.2
The Z.AI direct-supply GLM-5.2 model is now available. User guideAugust 6, 2026
wan3.0-video
Wan 3.0 All-in-One video generation model, now available for invite-only testing. Supports text-to-video, image-to-video (first-frame / first+last-frame), and reference-to-video in a single unified model. API referenceAugust 4, 2026
qwen-image-3.0-pro, qwen-image-3.0
The Qwen-Image-3.0-Pro model supports long-text input and dense image-in-image layouts, enabling precise one-shot generation of complex layouts such as newspapers, storyboards, menus, and exam papers. It features 10-pixel fine-grained text rendering, photographic-level detail reproduction of micro-expressions, pores and hair strands, and supports 12 languages, multiple fonts, and high-fidelity rendering of mainstream web and game interfaces. qwen-image-3.0 is the standard variant, balancing quality and speed.August 3, 2026
qwen3.8-max
The Qwen3.8 native vision-language Max series model is Qwen's most capable flagship model to date, featuring 2.4 trillion parameters with a Mixture-of-Experts architecture. It delivers outstanding performance comparable to the current state-of-the-art models, with significant improvements over the 3.7 series. Supports hybrid thinking mode (enabled by default) and a 1M-token context window. See Text generation models.August 1, 2026
deepseek-v4-flash-0731
A highly efficient, lightweight MoE model with 284 billion parameters in total and 13 billion activated parameters, natively supporting context windows of up to one million tokens. It offers fast inference speed, low latency, and cost-effective invocation, delivering well-balanced overall performance. Designed for high-concurrency, lightweight workloads, it is ideally suited for common, essential use cases such as everyday dialogue, content creation, basic RAG applications, and batch text processing. User guideJuly 30, 2026
qwen-audio-3.0-asr-flash-streaming, qwen-audio-3.0-asr-flash-filetrans, qwen-audio-3.0-asr-flash
Added the Qwen-Audio-3.0-ASR-Flash-Streaming (real-time), Qwen-Audio-3.0-ASR-Flash-Filetrans (non-real-time), and Qwen-Audio-3.0-ASR-Flash (non-real-time) speech recognition models:- Dialect support: Supports the seven major Chinese dialect groups (Mandarin, Wu, Xiang, Gan, Hakka, Min, and Yue) and more than 20 regional accents.
- Classical poetry optimization: Improves recognition accuracy for classical Chinese poetry, making it suitable for education, culture, and audiobook scenarios.
- Text optimization: Enhances punctuation prediction and text normalization, automatically converting numbers, dates, and monetary amounts to standard formats.
- Multilingual expansion: Supports 30 languages, including Chinese, English, Japanese, and Korean.
- Hotwords and context: Supports hotwords (precompiled and on-the-fly) and context input to improve recognition accuracy for domain-specific terms.
July 25, 2026
qwen3.7-flash, qwen3.7-flash-2026-07-15
The Qwen3.7 native vision-language Flash model series delivers a comprehensive upgrade over 3.6-Flash in multimodal understanding and agent execution. This model particularly excels in enhanced multimodal foundations with stronger universal object recognition, further improved real-world perception and spatial intelligence, significantly upgraded multimodal agent capabilities for Search Agent and CI Agent scenarios with more stable end-to-end task execution, as well as optimized multimodal coding for a smoother vibe coding experience.July 14, 2026
qwen-audio-3.0-tts-plus, qwen-audio-3.0-tts-flash
The Qwen-Audio-TTS speech synthesis model is now available. It adds support for more minority languages and Chinese dialects, with enhanced instruction following and fine-grained tag control. Audio quality and expressiveness are significantly improved. The Plus edition targets high-quality professional scenarios, while the Flash edition targets low-latency real-time interaction. User guideJuly 1, 2026
wan2.7-t2v-2026-06-12
A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same as wan2.7-t2v. API referencewan2.7-r2v-2026-06-12
A snapshot version of the Wan 2.7 reference-to-video model. It supports subject referencing and voice customization, and can generate a scripted video directly from a single, multi-panel storyboard image. API referenceJune 26, 2026
glm-5.2
The Zhipu GLM-5.2 model is the latest flagship designed for long-horizon tasks, supporting an ultra-long 1M context window. It features powerful logical reasoning, long-text comprehension, and code generation capabilities, balancing performance with inference efficiency. It excels across multi-task benchmarks and is well-suited for intelligent interaction, enterprise applications, and development assistance.kimi-k2.7-code
Kimi K2.7 Code model is now available in the Singapore region. An agent-centric coding model optimized for long-range software engineering tasks. Supports thinking mode only.June 25, 2026
qwen-image-2.0-pro-2026-06-22
The latest snapshot of the Qwen-Image-2.0 series, unifying image generation and editing. Compared with the April 22 snapshot, this model features enhanced text rendering with support for up to 1k token instructions, more refined photorealistic quality and scene detail, and stronger semantic adherence.June 22, 2026
HappyHorse 1.1
happyhorse-1.1-t2v
HappyHorse 1.1 text-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.1-i2v
HappyHorse 1.1 image-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.1-r2v
HappyHorse 1.1 reference-to-video model. Supports multiple reference images as input to generate audio videos with 3--15 seconds duration at 720P/1080P resolution. API referenceJune 17, 2026
fun-asr-flash-2026-06-15
Fun-ASR snapshot (June 2026 update). Adds full support for the seven major Chinese dialect systems (Mandarin, Wu, Xiang, Gan, Hakka, Min, Yue) and 20+ regional Mandarin accents, with specific optimizations for classical Chinese poetry recognition. Improves punctuation prediction and text normalization for numbers, dates, and amounts. Language support expanded to 30 languages. Supports contextualization and can transcribe audio up to 5 minutes.June 10, 2026
qwen3.7-max-2026-06-08
The Max model, the largest and most capable in the Qwen3.7 series, has added visual-modal understanding compared to the May 20 snapshot, enabling it to perceive real-world scenes and supporting multimodal interactive hybrid agent capabilities.June 1, 2026
qwen3.7-plus, qwen3.7-plus-2026-05-26
Qwen 3.7 Plus series builds upon strong text capabilities with comprehensively upgraded vision-language abilities while maintaining full agentic capabilities in coding, tool use, and productivity workflows. Its core differentiator is multimodal interactive hybrid agent capabilities — perceiving real-world scenarios, reading screens and operating GUIs, generating code based on visual references, and end-to-end navigation of mobile applications.May 27, 2026
glm-5.1
The Zhipu GLM-5.1 model is designed for long-horizon tasks and supports a 200K context window with a maximum output of 128K tokens. With strong logical reasoning, long-text comprehension, and code generation capabilities, it delivers excellent results across multiple benchmarks and is well suited for intelligent interaction, enterprise applications, and development assistance.May 25, 2026
qwen3.7-max-preview, qwen3.7-max-2026-05-17
Qwen Max series model snapshots. Text-only input, thinking mode enabled by default.May 22, 2026
qwen3.5-livetranslate-flash-realtime, qwen3.5-livetranslate-flash-realtime-2026-05-19
High-precision, real-time multilingual audio and video translation model built on Qwen3.5-Omni. Supports 60 languages (29 with audio + text output, 31 with text-only output), voice cloning for speaker-preserving translation, and visual context from video streams to improve accuracy. Upgrades from qwen3-livetranslate-flash-realtime with broader language coverage and lower latency. User guide | API referenceMay 21, 2026
qwen3.7-max, qwen3.7-max-2026-05-20
Next-generation flagship in the Qwen Max series. Text-only input, thinking mode enabled by default, supports explicit cache. Excels at coding, office & productivity, and long-horizon autonomous execution.May 11, 2026
deepseek-v4-pro, deepseek-v4-flash
DeepSeek V4 series. deepseek-v4-pro is a large-scale MoE model with strong general reasoning. deepseek-v4-flash is a lightweight, cost-effective model optimized for speed. Both support function calling, context cache, and thinking mode (enabled by default) with 1M context and a shared 384k output budget.April 27, 2026
HappyHorse video series
happyhorse-1.0-t2v
HappyHorse text-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-i2v
HappyHorse image-to-video model. Supports audio video generation with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-r2v
HappyHorse reference-to-video model. Supports multiple reference images as input to generate audio videos with 3--15 seconds duration at 720P/1080P resolution. API referencehappyhorse-1.0-video-edit
HappyHorse video editing model. Supports video editing and processing. API referenceApril 26, 2026
wan2.7-t2v-2026-04-25
A snapshot version of the Wan 2.7 text-to-video model. The model capabilities are the same aswan2.7-t2v. API referencewan2.7-i2v-2026-04-25
A snapshot version of the Wan 2.7 image-to-video model. The model capabilities are the same aswan2.7-i2v. API referenceApril 23, 2026
qwen3.5-plus-2026-04-20
New Qwen3.5-Plus snapshot with significantly improved agentic coding and faster inference compared to the February 15 snapshot. Knowledge, reasoning, and long-context capabilities remain strong — ideal for coding agents, production workflows, and high-throughput scenarios.qwen3.6-27b
Qwen3.6 27B dense open-source vision-language model. Enhanced agentic coding and STEM reasoning over Qwen3.5-27B. Significant improvements in spatial intelligence, object localization and detection. Steady gains in video understanding, document OCR, and vision agent capabilities. Supports function calling and structured output but not built-in tools.qwen-image-2.0-pro-2026-04-22
New Qwen-Image-2.0-Pro snapshot with unified image generation and editing. Compared to the March 3 snapshot, this model delivers noticeable improvements in visual quality — especially texture detail, lighting, and materials. Supports multilingual in-image text rendering and more balanced artistic style expression.April 20, 2026
qwen3.6-max-preview
The largest closed-source model in the Qwen3.6 series with improved coding and more efficient Agent execution. Supports text-only input, thinking mode (enabled by default), explicit cache, and Function Calling.April 16, 2026
qwen3.6-flash, qwen3.6-flash-2026-04-16, qwen3.6-35b-a3b
Qwen3.6 native vision-language Flash series with significant overall improvements over Qwen3.5-Flash. Enhanced agent programming capabilities with major benchmark gains over the previous generation, stronger math and code reasoning, and improved spatial intelligence — particularly object localization and detection.fun-asr, fun-asr-2025-11-07
Fun-ASR real-time speech recognition upgrades: dialect support covering seven major Chinese dialect groups and 20+ regional accents, improved classical poetry recognition, enhanced punctuation prediction and text normalization (numbers, dates, monetary values), and multilingual support for 30 languages.April 3, 2026
Wan 2.7 video series
wan2.7-videoedit
Instruction-based video editing and migration. Supports content replacement with reference images and replicating actions, effects, and camera movements.wan2.7-i2v
Multimodal input (text, image, audio, video) for first-frame, start-and-end-frame, and video continuation tasks.wan2.7-t2v
New resolution options and custom aspect ratios for different scenarios and platforms.wan2.7-r2v
Entity reference, voice customization, and playbook-based video generation from a single storyboard.April 2, 2026
qwen3.6-plus, qwen3.6-plus-2026-04-02
Upgraded coding (Agentic Coding, frontend, Vibe Coding), multimodal (object recognition, OCR, localization), and general reasoning. Fixes known issues from Qwen-3.5-Plus.April 1, 2026
wan2.7-image-pro, wan2.7-image
Text-to-image, image-to-image, image editing, multi-image reference, and interactive editing. Pro series supports 4K output.March 31, 2026