Skip to main content
Third-party models

DeepSeek

Call DeepSeek models through the OpenAI-compatible API or DashScope SDK on QwenCloud.

This guide shows how to call DeepSeek models via the OpenAI-compatible API or DashScope SDK.
The models deepseek-v3, deepseek-v3.1, deepseek-v3.2, deepseek-v3.2-exp, deepseek-r1, deepseek-r1-0528, and deepseek-r1-distill-qwen-7b/14b/32b will be deprecated on October 10, 2026. Migrate to qwen3.7-plus, qwen3.8-max, qwen3.8-flash, qwen3.7-max, or qwen3.6-flash.

Quick start

deepseek-v4-pro-0813 is the flagship model in the DeepSeek series with 1.6T total parameters and 49B activated parameters. It natively supports context windows of up to 1 million tokens and delivers top-tier performance across coding, math, and general tasks. You can use the enable_thinking parameter to switch between thinking and non-thinking modes. The following example calls deepseek-v4-pro-0813 in thinking mode. The newest addition, deepseek-v4.1-flash, is a lightweight flagship with 552B total parameters and only 8B activated for input. It also accepts image input. Calling it works the same way as the examples below, except that its DashScope endpoint differs — see Multimodal examples. Before you begin, get an API key and set it as an environment variable. If you call the model through an SDK, install the OpenAI or DashScope SDK.
  • OpenAI compatible
  • Anthropic compatible
  • DashScope
The enable_thinking parameter is not part of the standard OpenAI API. In the OpenAI Python SDK, pass it through extra_body. In the Node.js SDK, pass it as a top-level parameter. The reasoning_effort parameter is a standard OpenAI parameter that you can pass directly as a top-level parameter.
  • Python
  • Node.js
  • curl
Example code
from openai import OpenAI
import os

client = OpenAI(
  api_key=os.getenv("DASHSCOPE_API_KEY"),
  base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)

messages = [{"role": "user", "content": "Who are you?"}]
completion = client.chat.completions.create(
  model="deepseek-v4-pro-0813",
  messages=messages,
  extra_body={"enable_thinking": True},
  stream=True,
  stream_options={"include_usage": True},
)

reasoning_content = ""
answer_content = ""
is_answering = False
print("\n" + "=" * 20 + "Thinking process" + "=" * 20 + "\n")

for chunk in completion:
  if not chunk.choices:
    print("\n" + "=" * 20 + "Token usage" + "=" * 20 + "\n")
    print(chunk.usage)
    continue

  delta = chunk.choices[0].delta

  if hasattr(delta, "reasoning_content") and delta.reasoning_content is not None:
    if not is_answering:
      print(delta.reasoning_content, end="", flush=True)
    reasoning_content += delta.reasoning_content

  if hasattr(delta, "content") and delta.content:
    if not is_answering:
      print("\n" + "=" * 20 + "Full response" + "=" * 20 + "\n")
      is_answering = True
    print(delta.content, end="", flush=True)
    answer_content += delta.content

Reasoning effort

deepseek-v4.1-flash, deepseek-v4-pro-0813, deepseek-v4-pro, deepseek-v4-flash, and deepseek-v4-flash-0731 have thinking mode enabled by default. You can use the reasoning_effort parameter to control reasoning intensity. Valid values: low (supported only by deepseek-v4.1-flash, deepseek-v4-flash-0731, and deepseek-v4-pro-0813), high, and max. The default value is high.
For deepseek-v4-pro and deepseek-v4-flash, low and medium produce the same behavior as high, and xhigh produces the same behavior as max. For deepseek-v4.1-flash, minimal is mapped to low, medium and xhigh are mapped to high, and ultra is mapped to max.
  • OpenAI compatible
  • DashScope
  • Python
  • Node.js
  • curl
from openai import OpenAI
import os

client = OpenAI(
  api_key=os.getenv("DASHSCOPE_API_KEY"),
  base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
  model="deepseek-v4-pro-0813",
  messages=[{"role": "user", "content": "Which is larger, 9.9 or 9.11?"}],
  reasoning_effort="high",
)
print(completion.choices[0].message.content)

Responses API

deepseek-v4.1-flash, deepseek-v4-pro-0813, deepseek-v4-flash, deepseek-v4-flash-0731, and deepseek-v4-pro support calls through the OpenAI-compatible Responses API. When calling the Responses API, you can add the web_search (Web search), web_extractor (Web extractor), and code_interpreter (Code Interpreter) tools to the tools parameter.
  • Python
  • Node.js
  • curl
from openai import OpenAI
import os

client = OpenAI(
  # If the environment variable is not configured, replace the following line with your API key: api_key="sk-xxx"
  api_key=os.getenv("DASHSCOPE_API_KEY"),
  base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)
response = client.responses.create(
  model="deepseek-v4-flash",
  input="Hello! Please introduce yourself in one sentence.",
  # Optional: enable the web search, web extractor, and code interpreter tools
  tools=[
    {"type": "web_search"},
    {"type": "web_extractor"},
    {"type": "code_interpreter"},
  ],
)
# Get the model response
print(response.output_text)

Multimodal examples

deepseek-v4.1-flash has native visual understanding and accepts both text and image input, returning text. All other DeepSeek models accept text input only. Through the OpenAI-compatible API, pass images as image_url content blocks. Through DashScope, use the multimodal-generation endpoint with MultiModalConversation and pass images in the image field.
  • OpenAI compatible
  • DashScope
  • Python
  • Node.js
  • curl
from openai import OpenAI
import os

client = OpenAI(
  # If the environment variable is not configured, replace with your API key: api_key="sk-xxx"
  api_key=os.getenv("DASHSCOPE_API_KEY"),
  base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)

# Single image
completion = client.chat.completions.create(
  model="deepseek-v4.1-flash",
  messages=[
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {
          "type": "image_url",
          "image_url": {"url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241022/emyrja/dog_and_girl.jpeg"}
        }
      ]
    }
  ],
)
print(completion.choices[0].message.content)

# Multiple images (uncomment to use)
# completion = client.chat.completions.create(
#     model="deepseek-v4.1-flash",
#     messages=[
#         {
#             "role": "user",
#             "content": [
#                 {"type": "text", "text": "What is in these images?"},
#                 {
#                     "type": "image_url",
#                     "image_url": {"url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20241022/emyrja/dog_and_girl.jpeg"}
#                 },
#                 {
#                     "type": "image_url",
#                     "image_url": {"url": "https://dashscope.oss-cn-beijing.aliyuncs.com/images/tiger.png"}
#                 }
#             ]
#         }
#     ],
# )
# print(completion.choices[0].message.content)

Other features

ModelMulti-turnFunction callingWeb searchContext cacheStructured output
deepseek-v4.1-flash✓✓✓Implicit only✓
deepseek-v4-pro-0813✓✓✓Implicit only✓
deepseek-v4-pro✓✓✓✓—
deepseek-v4-flash✓✓✓Implicit only—
deepseek-v4-flash-0731✓✓✓Implicit only—
deepseek-v3.2✓✓✓✓—

Parameter defaults

Modeltemperaturetop_prepetition_penaltypresence_penaltymax_tokensthinking_budget
deepseek-v4.1-flash1.00.95--393,216 shared393,216 shared
deepseek-v4-pro-08131.00.95--393,216 shared393,216 shared
deepseek-v4-pro1.00.95--393,216 shared393,216 shared
deepseek-v4-flash1.00.95--393,216 shared393,216 shared
deepseek-v4-flash-07311.00.95--393,216 shared393,216 shared
deepseek-v3.21.00.95--65,53632,768
  • A hyphen (-) indicates that the parameter is not supported.
  • The deepseek-r1, deepseek-r1-0528, and distilled models do not support overriding their default parameter values.
  • For parameter descriptions, see the OpenAI-compatible Chat API.