vllm_omni.protocol.realtime.formats ¶
Realtime audio-format negotiation: what a client may declare and how it is read.
A Realtime client declares audio formats in three places --- the session object (session.audio.input.format and the older flat input_audio_format), a response.create override, and a conversation.item.create audio part --- and may spell one format several ways (audio/pcm, pcm16, s16le). This module is the single place that normalizes those spellings and says which are supported. Folding a session's declaration into the per-connection append defaults is the job of session.py, which builds on this module.
Everything here is a pure function of wire payloads: no session state, no model, no engine.
REALTIME_INPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_INPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"pcm_f32le",
"g711_ulaw",
"g711_alaw",
}
REALTIME_OUTPUT_AUDIO_FORMATS module-attribute ¶
REALTIME_OUTPUT_AUDIO_FORMATS = {
"pcm16",
"pcm_s16le",
"s16le",
"wav",
"pcm",
"g711_ulaw",
"g711_alaw",
}
parse_realtime_audio_format ¶
realtime_audio_format_object ¶
realtime_audio_format_object(
fmt: object, *, sample_rate_hz: int | None = None
) -> dict[str, object]
validate_conversation_item_audio_formats ¶
validate_realtime_response_audio_formats ¶
validate_realtime_session_audio_formats ¶
validate_realtime_session_audio_formats(
session_payload: Mapping[str, object],
*,
input_audio_formats: Collection[str] | None = None,
output_audio_formats: Collection[str] | None = None,
) -> str | None
Reject a session object that declares an audio format we cannot serve.
The format sets default to everything the codec can decode; a consumer that serves a narrower set passes its own (see vllm_omni.protocol.realtime.capabilities).