Full-Duplex WebSocket API¶
vLLM-Omni provides a full-duplex runtime for models that can continue receiving speech while producing speech. It adds persistent session state, model-specific turn policy, overlap handling, playback acknowledgement, and optional session resume.
Full duplex is distinct from the turn-based Realtime Audio API.
This page is the endpoint overview. For the complete wire contract, the vllm_omni.clients.duplex.DuplexClient Python library, and the per-model capability gates, see the Realtime Duplex API guide.
Choose an Endpoint¶
| Endpoint | Protocol | Recommended use |
|---|---|---|
WS /v1/realtime?duplex=1 | OpenAI Realtime-style events (the normative contract) | Applications and browser clients |
WS /v1/duplex | Alias of /v1/realtime?duplex=1 (same protocol, same handler) | Clients that prefer a dedicated path |
Python: DuplexOmni / InlineDuplexClient | Typed commands and events in-process | Embedding the model without a server |
Enable Full Duplex¶
A model is served full duplex when its registered pipeline declares a duplex_plugin (the model's DuplexModelPlugin) and its deploy configuration sets session_mode: duplex. vllm serve --omni constructs DuplexOmni instead of AsyncOmni. It serves /v1/realtime?duplex=1 (and its alias /v1/duplex), POST /v1/chat/completions, /v1/models and /health; every other turn-based HTTP route (speech, batch, embeddings, video, ...) answers "not available". Set session_mode: turn to select the ordinary online serving stack instead. The model's supported tasks and endpoint restrictions still apply.
/v1/chat/completions is the ordinary chat service running on the duplex engine: a request is a turn-based generation on the same stages, served alongside the live websocket sessions, with the request options the chat service supports. It is wired only when the model's plugin declares DuplexCapabilities.supports_chat_completions; on any other duplex model the route reports "not available". A model that should not serve the route at all lists it in the deploy configuration's endpoint_restrictions.
The deploy configuration of such a model must agree:
A duplex-capable model must explicitly set session_mode to duplex or turn in its deploy configuration, possibly through base_config inheritance. Missing or invalid values fail at startup. The default MiniCPM-o deployment continues to use duplex mode.
To run MiniCPM-o 4.5 with the turn-based engine:
vllm serve openbmb/MiniCPM-o-4_5 --omni \
--deploy-config vllm_omni/deploy/minicpmo_4_5_turn.yaml \
--trust-remote-code \
--port 8091
This profile inherits the default model and stage settings and overrides only session_mode. It supports ordinary HTTP requests without creating a duplex session handler. Mode selection happens at startup; a WebSocket query parameter does not switch engines. The Python Omni / AsyncOmni APIs are unchanged.
Warning
On a deployment that is not duplex, WS /v1/duplex fails with Duplex API is not available and /v1/realtime?duplex=1 falls back to the ordinary turn-based realtime handler. Confirm that session.created.session.capabilities is present before treating the connection as full duplex.
MiniCPM-o 4.5 (vllm_omni/deploy/minicpmo_4_5.yaml) is the only model served over this endpoint today. PersonaPlex and Nemotron VoiceChat still carry their pre-framework duplex code: their pipelines declare no duplex_plugin, so they run turn-based until the follow-up PRs port them to the plugin contract (RFC vllm-omni#7181).
JoyVL is a separate HTTP interaction orchestrator and does not use these WebSocket endpoints. See Standalone Experimental Servers.
MiniCPM-o Quick Start¶
Start the duplex deployment:
vllm serve openbmb/MiniCPM-o-4_5 --omni \
--deploy-config vllm_omni/deploy/minicpmo_4_5.yaml \
--trust-remote-code \
--port 8091
Stream a mono, PCM16, 16 kHz WAV file with the provided client:
python examples/online_serving/minicpmo/realtime_duplex_demo.py \
--url 'ws://localhost:8091/v1/realtime?duplex=1' \
--model openbmb/MiniCPM-o-4_5 \
--input-wav input_16k_mono.wav \
--ref-audio reference_voice.wav \
--output-dir /tmp/minicpmo-duplex
Realtime Event Lifecycle¶
A typical /v1/realtime?duplex=1 session follows this lifecycle:
- Send
session.updatewith the model, modalities, audio formats, and session options. - Wait for
session.created; inspectsession.capabilitiesinstead of assuming every model supports the same controls. - Send
input_audio_buffer.appendevents while microphone audio arrives. - Send
input_audio_buffer.commitat a user-turn boundary when required by the model policy. - Consume
response.created, transcript deltas,response.output_audio.delta, andresponse.doneorresponse.listenevents. - Send
playback.ackafter audio has been played when the session advertises playback acknowledgement support. - Send
session.closeand wait forsession.closed.
Unlike the turn-based realtime endpoint, input may continue while a response is active. The server can emit overlap.decision to describe whether input was deferred, treated as a short acknowledgement, or used to interrupt output.
Capabilities and Model Differences¶
The session.created payload includes capability fields such as supports_barge_in, supports_playback_ack, supports_multi_session, supports_session_resume, and chunk_period_ms. Treat this payload as the runtime contract and branch on the flags, never on the model name: a model that supports native overlapping speech may still advertise supports_barge_in=false when destructive output interruption and model-state rewind are not validated for it. Capacity and session-resume behavior also depend on the selected deployment configuration. The per-model table lives in the Realtime Duplex API guide.
Python API¶
vllm_omni.entrypoints.duplex_omni.DuplexOmni runs the same engine-resident sessions in-process: open_session() returns a DuplexSessionHandle whose submit() takes typed DuplexCommand objects and whose events() yields typed DuplexEvent objects (each with a to_realtime() wire rendering). vllm_omni.clients.inline_duplex.InlineDuplexClient exposes that handle behind the DuplexClient API. See the Realtime Duplex API guide.
See the MiniCPM-o example and the full-duplex runtime design for model-specific validation and architecture details.