Skip to content

vllm_omni.engine.duplex.turn_detection

Server-side turn detection (turn_detection: server_vad) for the session runner.

Moved from the Realtime input translator: the runner owns one ServerTurnDetector per session, feeds each appended chunk to it and emits input_audio_buffer.speech_started / speech_stopped from the result. The two-phase session.update commit (prepare / commit / reject) that used to live on the translator is expressed by :class:PendingTurnDetectionUpdate.

PendingTurnDetectionUpdate dataclass

A session.update turn-detection change awaiting the engine ACK (two-phase).

config instance-attribute

config: TurnDetectionConfig | None

detector instance-attribute

detector: ServerTurnDetector | None

commit

commit(
    current: ServerTurnDetector | None,
) -> tuple[
    TurnDetectionConfig | None, ServerTurnDetector | None
]

Adopt the staged detector; the replaced one is reset. Returns the new (config, detector).

prepare classmethod

prepare(
    session_payload: dict[str, object],
    *,
    backend_provider: SileroVADBackendProvider
    | None = None,
) -> PendingTurnDetectionUpdate | None

Normalize the payload in place and stage the new detector; None when not configured.

reject

reject() -> None

ServerTurnDetector

One Silero-backed streaming detector per session.

config instance-attribute

config = config

vad property

process

process(
    base64_audio: str,
    *,
    fmt: str,
    sample_rate_hz: int | None,
    audio_end_ms: int | None = None,
) -> TurnDetectionResult

Score one appended chunk. Raises ServerVADUnavailableError / ValueError like the VAD.

reset

reset() -> None

ServerVADUnavailableError

Bases: RuntimeError

SileroVADBackendProvider

Resolve, verify and load one detector backend per engine process.

ONNX is preferred and shared; the torch package is the fallback and is built per call because its state is internal.

model_path instance-attribute

model_path = model_path

get

The shared ONNX backend, or a per-session torch one when it cannot be built.

An explicitly configured server_vad_model_path never falls back: the operator named that artifact, so a missing file or a missing ONNX Runtime is their error, not something to paper over with a different model.

TurnDetectionConfig dataclass

Normalized server_vad configuration (all fields filled with defaults).

create_response class-attribute instance-attribute

create_response: bool = True

interrupt_response class-attribute instance-attribute

interrupt_response: bool = True

min_speech_duration_ms class-attribute instance-attribute

min_speech_duration_ms: int = 96

overlap_policy property

overlap_policy: str

prefix_padding_ms class-attribute instance-attribute

prefix_padding_ms: int = 300

silence_duration_ms class-attribute instance-attribute

silence_duration_ms: int = 500

threshold class-attribute instance-attribute

threshold: float = 0.5

build_detector

build_detector(
    backend_provider: SileroVADBackendProvider
    | None = None,
) -> ServerTurnDetector

from_realtime classmethod

from_realtime(
    turn_detection: Mapping[str, object],
) -> TurnDetectionConfig

TurnDetectionResult dataclass

audio_end_ms class-attribute instance-attribute

audio_end_ms: int | None = None

audio_start_ms class-attribute instance-attribute

audio_start_ms: int | None = None

create_response class-attribute instance-attribute

create_response: bool = True

is_speech instance-attribute

is_speech: bool

should_commit class-attribute instance-attribute

should_commit: bool = False

speech_active instance-attribute

speech_active: bool

speech_probability class-attribute instance-attribute

speech_probability: float = 0.0

speech_started class-attribute instance-attribute

speech_started: bool = False

speech_stopped class-attribute instance-attribute

speech_stopped: bool = False

as_payload

as_payload() -> dict[str, object]

apply_turn_detection_result

apply_turn_detection_result(
    payload: dict[str, object], result: TurnDetectionResult
) -> None

Merge a detector result onto an internal input_audio_buffer.append payload.

Mirrors what the old translator attached: is_speech from the detector, the vad hint block, and force_listen while speech is active.

configured_realtime_turn_detection

configured_realtime_turn_detection(
    session_payload: Mapping[str, object],
) -> tuple[str | None, object]

Return (field_path, value) for the turn detection setting present in the payload.

normalize_turn_detection_session_payload

normalize_turn_detection_session_payload(
    session_payload: dict[str, object],
) -> tuple[bool, TurnDetectionConfig | None]

Fill turn_detection defaults and derive overlap_policy in place.

Returns (configured, config) where configured is False when the payload does not mention turn detection at all (nothing changes), and config is None for turn_detection: null (model-owned turns, overlap_policy forced to listen_only).

validate_realtime_turn_detection

validate_realtime_turn_detection(
    session_payload: Mapping[str, object],
) -> str | None

Validate the turn_detection object of a Realtime session payload (None when valid).