Skip to content

vllm_omni.protocol.realtime.audio_input

Decoding one input_audio_buffer.append.

An append is the hottest client event on a Realtime connection and the one with the most wire surface: base64 audio in any accepted format, an optional rate, optional camera frames, and a set of optional hints (duration_ms, is_speech, a VAD probability, a transcript) that clients attach and servers are free to use or ignore.

:func:decode_audio_append turns that into :class:RealtimeAudioAppend, a value object: bytes at 16 kHz pcm_f32le, the surviving hints, and a speech verdict. It is a pure function of the event plus the session's :class:~vllm_omni.protocol.realtime.session.RealtimeInputDefaults, so any consumer decodes an append the same way. Building a consumer's own command object out of it is one constructor call --- see vllm_omni.engine.duplex.realtime_commands.build_append_audio, which wraps it for the duplex AppendAudio.

The speech verdict here is the client-declared one (explicit flags, a VAD probability the client sent, or an RMS floor as a last resort). It is not server-side VAD; that is a runtime concern and lives with whoever owns the session.

REALTIME_INPUT_HINT_KEYS module-attribute

REALTIME_INPUT_HINT_KEYS = (
    "duration_ms",
    "audio_duration_ms",
    "audio_start_ms",
    "audio_end_ms",
    "is_speech",
    "speech",
    "speech_probability",
    "vad",
    "overlap_action",
    "overlap",
    "force_barge_in",
    "force_listen",
    "text",
    "transcript",
)

RealtimeAudioAppend dataclass

One decoded input_audio_buffer.append.

audio is raw bytes in format at sample_rate_hz --- after :func:decode_audio_append that is 16 kHz pcm_f32le for every input format the codec converts.

audio instance-attribute

audio: bytes

audio_end_ms class-attribute instance-attribute

audio_end_ms: int | None = None

duration_ms class-attribute instance-attribute

duration_ms: int | None = None

event_id class-attribute instance-attribute

event_id: str | None = None

format instance-attribute

format: str

hints class-attribute instance-attribute

hints: dict[str, object] = field(default_factory=dict)

is_speech class-attribute instance-attribute

is_speech: bool | None = None

sample_rate_hz class-attribute instance-attribute

sample_rate_hz: int | None = None

video_frames class-attribute instance-attribute

video_frames: tuple[str, ...] = ()

copy_realtime_input_hints

copy_realtime_input_hints(
    source: Mapping[str, object], target: dict[str, object]
) -> None

decode_audio_append

decode_audio_append(
    event: Mapping[str, object],
    *,
    defaults: RealtimeInputDefaults,
    hints_source: Mapping[str, object] | None = None,
) -> RealtimeAudioAppend

Validate and pack one input append (audio and/or video frames).

Capability checks (required/optional modalities) run later on the session. Audio path converts to 16 kHz pcm_f32le.

hints_source is the enclosing payload when the audio arrives inside something larger than a bare append --- a conversation.item.create audio part, say --- so its hints apply unless the part overrides them.

Raises :class:RealtimeProtocolError for an unsupported format, undecodable audio or invalid camera frames.

input_explicitly_non_speech

input_explicitly_non_speech(
    event: Mapping[str, object],
) -> bool

input_looks_like_speech

input_looks_like_speech(
    event: Mapping[str, object],
    *,
    audio: object,
    fmt: object,
    overlap_silence_rms: float,
) -> bool