vllm_omni.engine.duplex.turn_detection ¶
Server-side turn detection (turn_detection: server_vad) for the session runner.
Moved from the Realtime input translator: the runner owns one ServerTurnDetector per session, feeds each appended chunk to it and emits input_audio_buffer.speech_started / speech_stopped from the result. The two-phase session.update commit (prepare / commit / reject) that used to live on the translator is expressed by :class:PendingTurnDetectionUpdate.
PendingTurnDetectionUpdate dataclass ¶
A session.update turn-detection change awaiting the engine ACK (two-phase).
commit ¶
commit(
current: ServerTurnDetector | None,
) -> tuple[
TurnDetectionConfig | None, ServerTurnDetector | None
]
Adopt the staged detector; the replaced one is reset. Returns the new (config, detector).
prepare classmethod ¶
prepare(
session_payload: dict[str, object],
*,
backend_provider: SileroVADBackendProvider
| None = None,
) -> PendingTurnDetectionUpdate | None
Normalize the payload in place and stage the new detector; None when not configured.
ServerTurnDetector ¶
ServerVADUnavailableError ¶
Bases: RuntimeError
SileroVADBackendProvider ¶
Resolve, verify and load one detector backend per engine process.
ONNX is preferred and shared; the torch package is the fallback and is built per call because its state is internal.
get ¶
get() -> SpeechDetectorBackend
The shared ONNX backend, or a per-session torch one when it cannot be built.
An explicitly configured server_vad_model_path never falls back: the operator named that artifact, so a missing file or a missing ONNX Runtime is their error, not something to paper over with a different model.
TurnDetectionConfig dataclass ¶
Normalized server_vad configuration (all fields filled with defaults).
build_detector ¶
build_detector(
backend_provider: SileroVADBackendProvider
| None = None,
) -> ServerTurnDetector
from_realtime classmethod ¶
from_realtime(
turn_detection: Mapping[str, object],
) -> TurnDetectionConfig
TurnDetectionResult dataclass ¶
apply_turn_detection_result ¶
apply_turn_detection_result(
payload: dict[str, object], result: TurnDetectionResult
) -> None
Merge a detector result onto an internal input_audio_buffer.append payload.
Mirrors what the old translator attached: is_speech from the detector, the vad hint block, and force_listen while speech is active.
configured_realtime_turn_detection ¶
configured_realtime_turn_detection(
session_payload: Mapping[str, object],
) -> tuple[str | None, object]
Return (field_path, value) for the turn detection setting present in the payload.
normalize_turn_detection_session_payload ¶
normalize_turn_detection_session_payload(
session_payload: dict[str, object],
) -> tuple[bool, TurnDetectionConfig | None]
Fill turn_detection defaults and derive overlap_policy in place.
Returns (configured, config) where configured is False when the payload does not mention turn detection at all (nothing changes), and config is None for turn_detection: null (model-owned turns, overlap_policy forced to listen_only).