Skip to main content
Tool calling

PDF understanding

PDF understanding enables the model to parse and comprehend PDF documents, extracting text and images for analysis. You can pass PDF files via URL or Base64 encoding.

PDF understanding enables the model to parse and comprehend PDF documents, extracting text and images from the document for analysis. You can pass a PDF file by URL or as Base64-encoded data through the OpenAI-compatible Chat Completions API or the DashScope API.
PDF understanding is not supported through the Responses API at this time. If you pass a PDF using the Responses API, the request still returns HTTP 200, but the file is not delivered to the model.

Supported models

qwen3.8-max, qwen3.8-max-0902, qwen3.8-flash

Quick start

Run the following code to send a PDF file to the model. Get a QwenCloud API key and configure it as an environment variable.
  • OpenAI compatible
  • DashScope
from openai import OpenAI
import os

client = OpenAI(
    # If the environment variable is not configured, replace it with your QwenCloud API key: api_key="sk-xxx"
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)

completion = client.chat.completions.create(
    model="qwen3.8-max",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "file",
                    "file": {
                        "file_url": "https://help-static-aliyun-doc.aliyuncs.com/file-manage-files/zh-CN/20260616/qmycjl/1506.02640v5.pdf"
                    }
                },
                {
                    "type": "text",
                    "text": "Summarize this PDF document"
                }
            ]
        }
    ],
    stream=True,
    stream_options={"include_usage": True},
    extra_body={"include_tool_usage": True}
)

for chunk in completion:
    if not chunk.choices:
        print(f"\nUsage: {chunk.usage}")
        continue
    delta = chunk.choices[0].delta
    if hasattr(delta, "content") and delta.content:
        print(delta.content, end="", flush=True)

Base64 input

If you cannot provide a file URL, you can also pass the PDF file as a Base64-encoded string. When using file_data, the filename field is required.
Base64 encoding increases the data size by about 1/3. For example, a 150 MB file becomes roughly 200 MB after encoding, which exceeds the request body size limit. For large files, pass the file by URL instead.
  • OpenAI compatible
  • DashScope
Python
import base64
from openai import OpenAI
import os

client = OpenAI(
    api_key=os.getenv("DASHSCOPE_API_KEY"),
    base_url="https://maas.qwencloudapi.com/compatible-mode/v1",
)

# Read and encode the PDF file
with open("report.pdf", "rb") as f:
    pdf_base64 = base64.b64encode(f.read()).decode("utf-8")

completion = client.chat.completions.create(
    model="qwen3.8-max",
    messages=[
        {
            "role": "user",
            "content": [
                {
                    "type": "file",
                    "file": {
                        "file_data": f"data:application/pdf;base64,{pdf_base64}",
                        "filename": "report.pdf"
                    }
                },
                {
                    "type": "text",
                    "text": "What are the key findings in this report?"
                }
            ]
        }
    ]
)

print(completion.choices[0].message.content)

Request parameters

The file input is specified as an element in the content array. The OpenAI-compatible protocol uses the type: "file" element, and the DashScope protocol uses an element containing file_url/file_data.
The URL field in the protocol only accepts a string. A list (array) of URLs is not supported.
OpenAI-compatible format:
ParameterTypeRequiredDescription
file_urlstringOne of the two requiredThe download URL of the PDF file. Mutually exclusive with file_data.
file_datastringOne of the two requiredBase64-encoded PDF input in the format data:application/pdf;base64,xxx. Mutually exclusive with file_url.
filenamestringConditionally requiredThe file name. Required when using file_data.
file_formatstringNoThe file format. Currently only pdf is supported. Defaults to pdf.

Limits

ItemLimit
Maximum file size150 MB
Maximum page count256 pages
PDF parsing may take longer than a regular text request. The first-token timeout is up to 300 seconds. We recommend using streaming output to get results in real time and avoid long waits.

Billing

Billing involves the following:
  • Model input tokens: Text and images parsed from the PDF are counted as input tokens, billed at the model's standard input token rate. For per-model input and output token prices, see Pricing.
  • Document parsing fee: Charged per page of the PDF document parsed, corresponding to the metering item document_parsing (pdf). $0.0033 per page.

View PDF page count

To view the parsed PDF page count in the response, get it from pdf_page_parser.count in the usage field. The two protocols behave slightly differently:
// Returned only when "include_tool_usage": true is passed at the top level of the request body.
// For the OpenAI Python SDK, pass it via extra_body={"include_tool_usage": True}.
// When using HTTP/curl directly, place it at the top level — do not nest it inside extra_body or stream_options.
{
    "usage": {
        "x_tools": {
            "pdf_page_parser": {
                "count": 10,
                "strategy": "normal"
            }
        }
    }
}

Error codes

If a call fails, see Error messages.