vllm_omni.engine.duplex.realtime_commands ¶
OpenAI Realtime client events -> :class:DuplexCommand (stateless).
This module is the duplex binding of the shared Realtime codec. Parsing, validating and decoding a client event is not duplex-specific and lives in vllm_omni.protocol.realtime, reached through vllm_omni.protocol.duplex; what lives here is the part that is: which :class:~vllm_omni.engine.duplex.commands.DuplexCommand a decoded event becomes, what the duplex engine can serve (:data:DUPLEX_REALTIME_CAPABILITIES), and the mapping between a Realtime output format and the duplex response_format vocabulary.
Everything that needs per-session state (input-buffer emptiness for commits, response-id fallbacks for cancels, conversation-item lookups, VAD) is resolved by the session runner through the helpers on :class:~vllm_omni.engine.duplex.realtime_events.RealtimeProjectionState; the commands produced here carry the raw client intent only.
The shared names re-exported at the bottom are a compatibility surface for existing importers. The canonical definitions are in vllm_omni.protocol.realtime and there is exactly one of each.
DUPLEX_REALTIME_CAPABILITIES module-attribute ¶
DUPLEX_REALTIME_CAPABILITIES = RealtimeProtocolCapabilities(
validate_turn_detection=_validate_duplex_turn_detection
)
REALTIME_INPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_INPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"pcm_f32le",
"g711_ulaw",
"g711_alaw",
}
REALTIME_INPUT_HINT_KEYS module-attribute ¶
REALTIME_INPUT_HINT_KEYS = (
"duration_ms",
"audio_duration_ms",
"audio_start_ms",
"audio_end_ms",
"is_speech",
"speech",
"speech_probability",
"vad",
"overlap_action",
"overlap",
"force_barge_in",
"force_listen",
"text",
"transcript",
)
REALTIME_OUTPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_OUTPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"wav",
"pcm",
"g711_ulaw",
"g711_alaw",
}
RealtimeInputDefaults dataclass ¶
Session-level wire defaults an append may omit (derived from the session object).
with_session_payload ¶
with_session_payload(
session_payload: Mapping[str, object],
) -> RealtimeInputDefaults
Return defaults updated from a Realtime session object (session.update).
apply_realtime_session_defaults ¶
apply_realtime_session_defaults(
defaults: RealtimeInputDefaults,
session_payload: Mapping[str, object],
) -> RealtimeInputDefaults
Derive the wire defaults a Realtime session object declares.
build_append_audio ¶
build_append_audio(
event: Mapping[str, object],
*,
defaults: RealtimeInputDefaults,
hints_source: Mapping[str, object] | None = None,
) -> AppendAudio
Decode one audio append (shared codec) and pack it as the duplex command.
copy_realtime_input_hints ¶
duplex_response_format ¶
A Realtime output format -> the duplex response_format vocabulary.
input_audio_transcription_config ¶
input_audio_transcription_config(
session_payload: Mapping[str, object],
) -> dict[str, object] | None
input_looks_like_speech ¶
input_looks_like_speech(
event: Mapping[str, object],
*,
audio: object,
fmt: object,
overlap_silence_rms: float,
) -> bool
json_safe_realtime_payload ¶
normalize_conversation_item ¶
parse_realtime_audio_format ¶
realtime_audio_format_object ¶
realtime_audio_format_object(
fmt: object, *, sample_rate_hz: int | None = None
) -> dict[str, object]
realtime_max_output_tokens ¶
Normalize Realtime max output tokens ("inf" -> None).
realtime_overlap_fields ¶
text_chars_for_audio_ms_from_marks ¶
text_chars_for_audio_ms_from_marks(
audio_end_ms: int,
text_len: int,
marks: list[object],
*,
final_ms: object | None = None,
) -> int
translate_realtime_command ¶
translate_realtime_command(
payload: Mapping[str, object],
*,
defaults: RealtimeInputDefaults | None = None,
) -> DuplexCommand
Map one OpenAI Realtime client event onto a :class:DuplexCommand.
Raises :class:DuplexCommandError for malformed or unsupported payloads. session.resume and session.event_ack are transport concerns and are rejected with code="unknown_event".
truncate_realtime_item_content ¶
truncate_realtime_item_content(
item: dict[str, object],
*,
content_index: int,
audio_end_ms: int,
) -> None
validate_conversation_item_audio_formats ¶
validate_realtime_item_truncate ¶
validate_realtime_item_truncate(
item: Mapping[str, object],
*,
content_index: int,
audio_end_ms: int,
) -> str | None
validate_realtime_response_audio_formats ¶
validate_realtime_session_audio_formats ¶
validate_realtime_session_audio_formats(
session_payload: Mapping[str, object],
*,
input_audio_formats: Collection[str] | None = None,
output_audio_formats: Collection[str] | None = None,
) -> str | None
Reject a session object that declares an audio format we cannot serve.
The format sets default to everything the codec can decode; a consumer that serves a narrower set passes its own (see vllm_omni.protocol.realtime.capabilities).
validate_realtime_video_frames ¶
Validate omni-duplex camera frames on input_audio_buffer.append.
Wire contract matches the official MiniCPM-o duplex loop: one base base64 JPEG per ~1 s audio chunk, optionally followed by that unit's stacked composite tiling the sub-frames captured inside it (at most 2 images either way). A caller-supplied max_slice_nums is rejected rather than silently ignored: slicing is Stage 0's decision here, and Stage 0 already applies the official HD suggestion for a stacked unit (max_slice_nums=[2, 1]). The wire simply does not let the client choose it.