Skip to main content
Text generation

Text generation models

Choose a model for AI agents, chatbots, document processing, and more.

Using OpenClaw or Claude Code?

qwen3.7-plus — balanced performance and cost, full tool support, 1M context for large codebases. For strongest reasoning, choose the flagship qwen3.8-max.

Migrate from closed-source models

Map your current GPT, Claude, or Gemini model to an equivalent QwenCloud model.
TierClosed-source examplesQwenCloud recommendation
Highest capabilityGPT-5.5, Claude Opus 4.7, Gemini 3.1 Proqwen3.8-max
BalancedGPT-5.4, Claude Sonnet 4.6, Gemini 3 Proqwen3.7-plus, deepseek-v4-pro-0813, glm-5.2
Lightweight & low-costGPT-5.4-mini, Claude Haiku 4.5, Gemini 3.1 Flashqwen3.8-flash, deepseek-v4.1-flash

For other applications

Chatbots, content generation, summarization, document processing — start with qwen3.7-plus, which balances performance, cost, and built-in tools with a 1M-token context window. To cut costs, switch to qwen3.8-flash, which offers similar capabilities at a lower price. For the strongest reasoning, use the flagship qwen3.8-max (1M context, higher cost).

Office productivity (non-coding)

For non-coding office tasks such as document drafting, email composition, meeting note summarization, and data analytics, start with qwen3.7-plus — it balances performance and cost with a 1M-token context window, Function Calling support, and built-in tools. To reduce costs, try qwen3.8-flash, which delivers near-flagship performance at a lower price with the same context length. For the strongest reasoning capability, choose the flagship qwen3.8-max (higher cost). For long-document processing such as reviewing multiple contracts, use qwen-long (10M-token context window).
TONGYI Lingma and Qoder are AI coding tools designed for software development. They are not intended for general office productivity tasks.

Context window

1M tokens is roughly 750,000 words or 10 novels.
  • Long documents or large codebases → qwen3.8-max / qwen3.7-max / qwen3.7-plus / qwen3.8-flash (1M)
  • Standard tasks → 128k–256k is plenty

Thinking mode

Step-by-step reasoning for multi-step math, debugging, architecture planning, or legal cross-referencing. Toggle with enable_thinking. All Qwen3+ models support it — most are hybrid, so you can switch per request.

Function calling + built-in tools

Let the model take actions: check weather, query a database, book a meeting.
  • Function calling (you define tools, model calls them): all general-purpose models
  • Built-in tools (web search, code execution — no setup): qwen3.8-max, qwen3.8-max-0902, qwen3.8-flash, qwen3.7-max, qwen3.7-plus, qwen3.7-flash, qwen3.6-flash, qwen3.5-plus, qwen3.5-flash, qwen3-max series only
ModelContextThinkingFunction callingBuilt-in toolsStructured output
qwen3.8-max1M✓✓✓✓
qwen3.8-flash1M✓✓✓✓
qwen3.7-plus1M✓✓✓✓
qwen3.6-flash1M✓✓✓✓
deepseek-v4.1-flash1M✓✓✓✓
deepseek-v4-pro-08131M✓✓✓✓
deepseek-v4-flash1M✓✓—✓
deepseek-v4-flash-07311M✓✓——
glm-5.2198k✓✓—✓

All models

Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen3.8-max1M128k256k✓✓✓
qwen3.8-max-09021M128k256k✓✓✓
qwen3.8-flash1M128k256k✓✓✓
qwen3.8-2.4t-a95b1M128k128k✓✓✓
Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen3.7-max1M64k256k✓✓—
qwen3.7-max-2026-06-081M64k256k✓✓—
qwen3.7-max-2026-05-201M64k256k✓✓—
qwen3.7-max-preview1M64k256k✓✓—
qwen3.7-max-2026-05-171M64k256k✓✓—
qwen3.7-plus1M64k256k✓✓✓
qwen3.7-plus-2026-05-261M64k256k✓✓✓
qwen3.7-flash1M64k256k✓✓✓
qwen3.7-flash-2026-07-151M64k256k✓✓✓
Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen3.6-max-preview256k64k128k✓—✓
qwen3.6-flash1M64k80k✓✓✓
qwen3.6-flash-2026-04-161M64k80k✓✓✓
qwen3.6-35b-a3b256k64k80k✓✓✓
qwen3.6-27b256k64k80k✓—✓
qwen3.6-plus1M64k80k✓✓✓
qwen3.6-plus-2026-04-021M64k80k✓✓✓
Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen3.5-plus1M64k80k✓✓✓
qwen3.5-plus-2026-04-201M64k80k✓✓✓
qwen3.5-plus-2026-02-151M64k80k✓✓✓
qwen3.5-flash1M64k80k✓✓✓
qwen3.5-flash-2026-02-231M64k80k✓✓✓
qwen3.5-397b-a17b256k64k80k✓✓✓
qwen3.5-122b-a10b256k64k80k✓✓✓
qwen3.5-27b256k64k80k✓✓✓
qwen3.5-35b-a3b256k64k80k✓✓✓

Translation

Model IDContextMax OutputFunction callingBuilt-in toolsStructured output
qwen-mt-plus16k8k———
qwen-mt-turbo16k8k———
qwen-mt-flash16k8k———
qwen-mt-lite16k8k———

Character roleplay

Model IDContextMax OutputFunction callingBuilt-in toolsStructured output
qwen-plus-character32k4k———
qwen-plus-character-ja8k4k———
qwen-flash-character8k4k———
Non-Qwen models available through the same API.
Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
deepseek-v4.1-flash1M384k **✓✓✓
deepseek-v4-pro-08131M384k **✓✓✓
deepseek-v4-pro1M384k **✓—✓
deepseek-v4-flash1M384k **✓—✓
deepseek-v4-flash-07311M384k **✓——
deepseek-v3.2128k64k32k✓——
* DeepSeek V4 models share a 384k total budget across output and thinking.
The following models are legacy and no longer recommended. Visit the Model Marketplace to view detailed model parameters, such as context window and billing.

Qwen3

Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen3-max256k64k80k✓✓✓
qwen3-max-2026-01-23256k64k80k✓✓✓
qwen3-max-preview256k64k80k✓✓✓
qwen3-max-2025-09-23256k64k—✓✓✓
qwen3-235b-a22b128k16k38k✓—✓
qwen3-235b-a22b-thinking-2507128k32k80k✓——
qwen3-235b-a22b-instruct-2507128k32k—✓—✓
qwen3-next-80b-a3b-thinking128k32k80k✓——
qwen3-next-80b-a3b-instruct128k32k—✓—✓
qwen3-32b128k16k38k✓—✓
qwen3-30b-a3b128k16k38k✓—✓
qwen3-30b-a3b-thinking-2507128k32k80k✓——
qwen3-30b-a3b-instruct-2507128k32k—✓—✓
qwen3-14b128k8k38k✓—✓
qwen3-8b128k8k38k✓—✓
qwen3-4b128k8k38k✓—✓
qwen3-1.7b32k8k30k✓—✓
qwen3-0.6b32k8k30k✓—✓

Qwen3-Coder

Model IDContextMax OutputFunction callingBuilt-in toolsStructured output
qwen3-coder-plus1M64k✓—✓
qwen3-coder-plus-2025-09-231M64k✓—✓
qwen3-coder-plus-2025-07-221M64k✓—✓
qwen3-coder-flash1M64k✓—✓
qwen3-coder-flash-2025-07-281M64k✓—✓
qwen3-coder-next256k64k✓—✓
qwen3-coder-480b-a35b-instruct256k64k✓—✓
qwen3-coder-30b-a3b-instruct256k64k✓—✓

Qwen2.5 (open source)

Model IDContextMax OutputFunction callingBuilt-in toolsStructured output
qwen2.5-omni-7b32k8k✓—✓
qwen2.5-vl-72b-instruct128k8k✓—✓
qwen2.5-vl-32b-instruct128k8k✓—✓
qwen2.5-vl-7b-instruct128k8k✓—✓
qwen2.5-vl-3b-instruct128k8k✓—✓
qwen2.5-72b-instruct32k8k✓—✓
qwen2.5-32b-instruct32k8k✓—✓
qwen2.5-14b-instruct32k8k✓—✓
qwen2.5-14b-instruct-1m1M8k✓—✓
qwen2.5-7b-instruct32k8k✓—✓
qwen2.5-7b-instruct-1m1M8k✓—✓

Legacy (qwen-plus/max/flash/turbo)

Model IDContextMax OutputThinking BudgetFunction callingBuilt-in toolsStructured output
qwen-plus1M32k80k✓—✓
qwen-plus-latest1M32k80k✓—✓
qwen-plus-2025-12-011M32k80k✓—✓
qwen-plus-2025-09-111M32k80k✓—✓
qwen-plus-2025-07-281M32k80k✓—✓
qwen-plus-2025-07-14128k16k80k✓—✓
qwen-plus-2025-04-28128k16k80k✓—✓
qwen-plus-2025-01-25128k8k—✓—✓
qwen-max32k8k—✓—✓
qwen-max-latest32k8k—✓—✓
qwen-max-2025-01-2532k8k—✓—✓
qwen-flash1M32k80k✓—✓
qwen-flash-2025-07-281M32k80k✓—✓
qwen-turbo128k16k38k✓—✓
qwen-turbo-latest128k16k38k✓—✓
qwen-turbo-2025-04-28128k16k38k✓—✓
qwen-turbo-2024-11-011M8k—✓—✓
qwq-plus128k8k32k———
qvq-max128k8k80k———
qvq-max-latest128k8k80k———
qvq-max-2025-03-25128k8k80k———
qwen-omni-turbo32k2k80k———
qwen-omni-turbo-latest32k2k80k———
qwen-omni-turbo-2025-03-2632k2k80k———

Learn more