WebSocket server reference
Server events for the Qwen-Omni-Realtime API, including tool calling (function calling) events.
Server error message.
First event after connection. Contains the default session configuration.
Sent after a successful
Sent in VAD mode when speech starts in the audio buffer.
Sent in VAD mode when speech ends in the audio buffer. The server also sends
Sent when the input audio buffer is committed.
Sent after the client sends
Sent when a conversation item is created.
When input audio transcription is enabled, this event is sent frequently while the user is speaking. It provides real-time intermediate transcription results. Concatenate
Sent after audio is buffered and transcribed. Transcription uses a separate model (
Sent when input audio transcription fails (if enabled). Separate from the
Sent when the model starts generating a response.
Sent after response generation completes. The
Sent when the output modality is text-only and the model generates a text chunk.
Sent when text-only output finishes generating.
Sent when the output modality includes audio and the model generates an audio chunk.
Sent when audio output finishes generating.
Sent when the output modality includes audio and the model generates a transcript chunk.
Sent when the audio transcript finishes generating.
When the model generates the argument string for a function call in a streaming manner, the server pushes this event for each new segment. Concatenate the
Indicates that the function call arguments have been fully generated. The
Sent when a new item is created during response generation. The item type can be
Sent when an output item is complete.
Sent when a new content part is added to an assistant message during response generation.
Sent when a content part in an assistant message finishes streaming.
This section describes WebSocket events and fields for
All three require string fields
Requires
All three require
These objects are types of the
Returned as
Each tool requires
When the model selects an MCP tool, this item appears in
Returned in
After a call reaches a terminal state, the item appears in
Successful call:
Failed call:
The error object requires string fields
Reference: Real-time multimodal.
error
Server error message.
Example
string
body
Unique event identifier.
string
body
Always
error.object
body
Error details.
session.created
First event after connection. Contains the default session configuration.
Example
string
body
Unique event identifier.
string
body
Always
session.created.object
body
Session configuration.
session.updated
Sent after a successful session.update request. On error, the server sends an error event instead.
Example
string
body
Unique event identifier.
string
body
Always
session.updated.object
body
Session configuration.
input_audio_buffer.speech_started
Sent in VAD mode when speech starts in the audio buffer.
May also fire each time audio is added to the buffer before speech is detected.
Example
string
body
Unique event identifier.
string
body
Always
input_audio_buffer.speech_started.integer
body
Milliseconds from the start of audio input to the first detected speech.
string
body
User message item ID, created when speech stops. This item appends user input to the conversation history for inference.
input_audio_buffer.speech_stopped
Sent in VAD mode when speech ends in the audio buffer. The server also sends conversation.item.created to create the user message item.
Example
string
body
Unique event identifier.
string
body
Always
input_audio_buffer.speech_stopped.integer
body
Milliseconds from session start to speech end.
string
body
User message item ID (will be created).
input_audio_buffer.committed
Sent when the input audio buffer is committed.
- In VAD mode, the buffer commits automatically when the user finishes speaking.
-
In manual mode, sent after the client sends
input_audio_buffer.commit.
Example
string
body
Unique event identifier.
string
body
Always
input_audio_buffer.committed.string
body
User message item ID (will be created).
input_audio_buffer.cleared
Sent after the client sends input_audio_buffer.clear.
Example
string
body
Unique event identifier.
string
body
Always
input_audio_buffer.cleared.conversation.item.created
Sent when a conversation item is created.
Example
string
body
Unique event identifier.
string
body
Always
conversation.item.created.object
body
Conversation item.
conversation.item.input_audio_transcription.delta
When input audio transcription is enabled, this event is sent frequently while the user is speaking. It provides real-time intermediate transcription results. Concatenate text + stash to get the most complete sentence preview at any point in time.
Example
string
body
Unique event identifier.
string
body
Always
conversation.item.input_audio_transcription.delta.string
body
The ID of the associated conversation item.
integer
body
The index of the content part that contains the audio.
string
body
The confirmed text prefix. This portion of the current sentence has been confirmed by the model and will not change.
string
body
The preliminary text suffix. This temporary draft follows the confirmed portion and may be revised by the model.
string
body
The detected language of the recognized audio.
string
body
The detected emotion of the recognized audio. Valid values:
neutral, happy, sad, angry, surprised, disgusted, fearful.Example: how text and stash fields work together
Example: how text and stash fields work together
Suppose the user says: "The weather is nice today, sunny and warm."
| Time | User speech | text | stash | Display (text + stash) |
|---|---|---|---|---|
| T1 | "The weather..." | "" | "The weather" | The weather |
| T2 | "...is nice..." | "" | "The weather is nice" | The weather is nice |
| T3 | "...today," | "The weather" | " is nice today," | The weather is nice today, |
| T4 | (brief pause) | "The weather is nice today, " | "" | The weather is nice today, |
| T5 | "sunny..." | "The weather is nice today, " | "sunny" | The weather is nice today, sunny |
| T6 | "...and warm." | "The weather is nice today, " | "sunny and warm." | The weather is nice today, sunny and warm. |
| T7 | (stops) | - | - | Use conversation.item.input_audio_transcription.completed as the final result. |
conversation.item.input_audio_transcription.completed
Sent after audio is buffered and transcribed. Transcription uses a separate model (qwen3-asr-flash-realtime).
The transcribed text may differ from text processed by Qwen-Omni-Realtime. Treat it as a reference.
Example
string
body
Unique event identifier.
string
body
Always
conversation.item.input_audio_transcription.completed.string
body
User message item ID.
integer
body
Fixed to 0.
string
body
Transcribed text.
conversation.item.input_audio_transcription.failed
Sent when input audio transcription fails (if enabled). Separate from the error event.
Example
string
body
Unique event identifier.
string
body
Always
conversation.item.input_audio_transcription.failed.string
body
User message item ID.
integer
body
Fixed to 0.
object
body
Error details.
response.created
Sent when the model starts generating a response.
Example
string
body
Unique event identifier.
string
body
Always
response.created.object
body
Response object.
response.done
Sent after response generation completes. The response object contains all output items except raw audio data.
Example
string
body
Unique event identifier.
string
body
Always
response.done.object
body
Response object.
response.text.delta
Sent when the output modality is text-only and the model generates a text chunk.
Example
string
body
Unique event identifier.
string
body
Always
response.text.delta.string
body
Incremental text chunk.
string
body
Response ID.
string
body
Message item ID. Use this to associate items from the same message.
integer
body
Output item index. Fixed to 0.
integer
body
Content part index. Fixed to 0.
response.text.done
Sent when text-only output finishes generating.
Also sent when the response is interrupted, incomplete, or canceled.
Example
string
body
Unique event identifier.
string
body
Always
response.text.done.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
integer
body
Content part index.
string
body
Complete text output.
response.audio.delta
Sent when the output modality includes audio and the model generates an audio chunk.
Example
string
body
Unique event identifier.
string
body
Always
response.audio.delta.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
integer
body
Content part index.
string
body
Base64-encoded audio chunk.
response.audio.done
Sent when audio output finishes generating.
Also sent when the response is interrupted, incomplete, or canceled.
Example
string
body
Unique event identifier.
string
body
Always
response.audio.done.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
integer
body
Content part index.
response.audio_transcript.delta
Sent when the output modality includes audio and the model generates a transcript chunk.
Example
string
body
Unique event identifier.
string
body
Always
response.audio_transcript.delta.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
integer
body
Content part index.
string
body
Incremental transcript text.
response.audio_transcript.done
Sent when the audio transcript finishes generating.
Example
string
body
Unique event identifier.
string
body
Always
response.audio_transcript.done.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
integer
body
Content part index.
string
body
Complete transcript.
response.function_call_arguments.delta
When the model generates the argument string for a function call in a streaming manner, the server pushes this event for each new segment. Concatenate the delta fields in order. The complete content is provided in the subsequent response.function_call_arguments.done event.
Example
string
body
Unique event identifier.
string
body
Always
response.function_call_arguments.delta.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
string
body
Unique ID for this function invocation. Consistent with the
done event in the same turn.string
body
New segment of the argument string. Concatenate segments in order.
response.function_call_arguments.done
Indicates that the function call arguments have been fully generated. The arguments field contains the complete argument string. After receiving this event, parse the arguments and call the local tool function. Use the complete arguments from this event, not the concatenated delta result.
Example
string
body
Unique event identifier.
string
body
Always
response.function_call_arguments.done.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index.
string
body
Unique ID for this function invocation.
string
body
Name of the function that was called.
string
body
Complete arguments for the function invocation, typically a JSON string.
response.output_item.added
Sent when a new item is created during response generation. The item type can be message or function_call; Qwen3.8-Omni-Flash-Realtime also supports mcp_call.
Example
string
body
Unique event identifier.
string
body
Always
response.output_item.added.string
body
Response ID.
integer
body
Output item index.
object
body
Output item.
response.output_item.done
Sent when an output item is complete.
Example
string
body
Unique event identifier.
string
body
Always
response.output_item.done.string
body
Response ID.
integer
body
Output item index.
object
body
Output item.
response.content_part.added
Sent when a new content part is added to an assistant message during response generation.
Example
string
body
Unique event identifier.
string
body
Always
response.content_part.added.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index. Fixed to 0.
integer
body
Content part index. Fixed to 0.
object
body
Content part.
response.content_part.done
Sent when a content part in an assistant message finishes streaming.
Example
string
body
Unique event identifier.
string
body
Always
response.content_part.done.string
body
Response ID.
string
body
Message item ID.
integer
body
Output item index. Fixed to 0.
integer
body
Content part index. Fixed to 0.
object
body
Content part.
MCP events
This section describes WebSocket events and fields for qwen3.8-omni-flash-realtime. Other base events and fields use the common protocol documented above. Required fields are required when their containing object is present. IDs are opaque strings; do not rely on their lengths, prefixes, or generation rules. arguments and output are strings whose contents require JSON parsing. MCP connections, discovery, and calls are subject to service quotas and timeouts.
mcp_list_tools.*
| Event type | When emitted |
|---|---|
mcp_list_tools.in_progress | Tool discovery starts for an MCP server. |
mcp_list_tools.completed | Discovery succeeds and allowed_tools filtering completes. |
mcp_list_tools.failed | Discovery fails, times out, or exceeds service limits. |
event_id (unique event ID), type (event type above), and item_id (associated mcp_list_tools item). in_progress can arrive before session.updated. For completed and failed, conversation.item.created containing the final tool list or error arrives first, followed by the status event with the same item_id.
response.mcp_call_arguments.delta
| Field | Type | Required | Meaning |
|---|---|---|---|
event_id | string | Yes | Unique event ID. |
type | string | Yes | response.mcp_call_arguments.delta |
response_id | string | Yes | Parent Response ID. |
item_id | string | Yes | Current mcp_call item ID. |
output_index | integer | Yes | Zero-based index in the parent Response output array. |
delta | string | Yes | Incremental JSON-string fragment; concatenate in event order. |
obfuscation | string | No | Optional obfuscation string; ignore when assembling and parsing arguments. |
response.mcp_call_arguments.done
Requires event_id, type, response_id, and item_id (strings), output_index (integer, same meaning as above), and arguments (string). type is response.mcp_call_arguments.done. arguments contains the complete JSON-encoded arguments and is authoritative.
response.mcp_call.*
| Event type | When emitted |
|---|---|
response.mcp_call.in_progress | Arguments are complete and approval has been granted or is unnecessary; the MCP call starts. |
response.mcp_call.completed | The MCP call succeeds. |
response.mcp_call.failed | Connection, protocol, or tool business error; approval rejection or timeout; call timeout; or cancellation. |
event_id, type, and item_id (strings), and output_index (integer). item_id identifies the current mcp_call; output_index is its index in the parent Response output array.
MCP conversation items
These objects are types of the item field in events, applicable to Qwen3.8-Omni-Flash-Realtime.
mcp_list_tools item
Returned as conversation.item.created.item.
| Field | Type | Required | Meaning |
|---|---|---|---|
item.id | string | Yes | Unique tool-list item ID, matching discovery event item_id. |
item.type | string | Yes | mcp_list_tools |
item.server_label | string | Yes | Configured MCP server label. |
item.tools | array of objects | Yes | Validated, filtered tool definitions; empty on failure. |
item.error | object | No | Discovery error; see MCP error below. |
name (string, original MCP tool name, 1-64 letters/digits/underscores/periods/hyphens) and input_schema (object, mapped from MCP inputSchema, with root type="object"). Optional fields are description (string from the server) and annotations (object, MCP ToolAnnotations).
Standard JSON Schema fields including properties, required, additionalProperties, $defs, oneOf, anyOf, and allOf are preserved and forwarded. The service does not perform full JSON Schema semantic validation; the MCP server validates the actual arguments.
Initial mcp_call item
When the model selects an MCP tool, this item appears in response.output_item.added.item and may also appear in conversation.item.created.item. All fields below are required strings.
| Field | Meaning |
|---|---|
item.id | Unique call item ID. |
item.object | realtime.item |
item.type | mcp_call |
item.status | Initially in_progress. |
item.call_id | Unique tool-call ID. |
item.server_label | Server actually used for this call. |
item.name | Original tool name. |
item.arguments | Usually empty initially; use arguments.done and the final item for complete arguments. |
mcp_approval_request item
Returned in conversation.item.created. The event requires string fields event_id, type (conversation.item.created), response_id (parent Response), item_id (approval ID, equal to item.id), and previous_item_id (the mcp_call item being reviewed), plus the item object.
All item fields are required strings: id (copy unchanged into the approval response's approval_request_id), type (mcp_approval_request), server_label, name, arguments (complete JSON-encoded arguments), and call_id (tool-call ID).
Final mcp_call item
After a call reaches a terminal state, the item appears in response.output_item.done.item and in the final parent response.done.response.output snapshot.
| Field | Type | Required | Meaning |
|---|---|---|---|
item.id | string | Yes | Unique call item ID. |
item.type | string | Yes | mcp_call |
item.status | string | Yes | completed or failed. |
item.call_id | string | Yes | Tool-call ID. |
item.server_label | string | Yes | Server actually used. |
item.name | string | Yes | Original tool name. |
item.arguments | string | Yes | Complete JSON-encoded arguments. |
item.output | string | No | Serialized MCP tools/call.result object; may also appear on some failures. |
item.error | object | No | Call error; see below. |
MCP error
The error object requires string fields type (fixed to tool_execution_error) and message (a client-safe error description that excludes sensitive upstream response bodies). If the MCP server returns isError=true, the final call has status="failed" and may include both the original output and a structured error.
See Qwen-Omni-Realtime.