vllm_omni.engine.duplex.plugin ¶
Model plugin contract for full-duplex models.
One DuplexModelPlugin subclass per model binds what used to be two dotted paths (the engine DuplexRuntimeExtension and the serving ServingRuntimeAdapter). Everything runs engine-side now, so the plugin is loaded once by DuplexOmniEngine and handed to DuplexOrchestrator / DuplexSessionManager.
EncodeAudio module-attribute ¶
DuplexDataPlane ¶
DuplexModelPlugin ¶
Bases: ABC
Everything vLLM-Omni needs to know about one full-duplex model.
Engine policy (sampling params, append planning, output decisions) and session policy (capabilities, runtime configuration, per-session state, data-plane projection) live on the same object so a mismatch between the two halves is impossible by construction.
private_runtime_config_keys class-attribute instance-attribute ¶
projects_intermediate_outputs class-attribute instance-attribute ¶
projects_intermediate_outputs: bool = False
silence_continuation_samples class-attribute instance-attribute ¶
silence_continuation_samples: int = 16000
commit_model_context ¶
Persist model-context history at a turn boundary. Default is a no-op.
This is not playback-ACK history. A model that keeps its own prompt transcript implements this; the session runner only decides when.
configure_sampling_params abstractmethod ¶
configure_sampling_params(
*,
runtime_config: dict[str, object],
defaults: tuple[object, ...],
) -> tuple[object, ...]
data_plane_context abstractmethod ¶
data_plane_context(
*,
epoch: int,
turn_id: int,
active_response_turn_id: int | None,
active_response_id: str | None,
auto_responds: bool,
response_format: str,
speed: float | None,
modalities: tuple[str, ...],
) -> object
decide_output abstractmethod ¶
decide_output(
*,
stage_id: int,
final_stage_id: int,
segment_finished: bool,
segment_token_ids: tuple[int, ...],
segment_output_metadata: dict[str, object],
output: object,
) -> DuplexOutputDecision | None
draining_stage_ids ¶
Output stages that may keep running after the next user turn starts.
Empty means a concurrent turn does not overlap a previous output stage. Shared lifecycle reads this instead of assuming a stage layout.
partial_stage_followup ¶
partial_stage_followup(
plan: PartialStageForward, req_state: object
) -> PartialStageForward | None
Optional second submit after plan has already been forwarded.
Default models have nothing to add. AURA uses this to queue the end sentinel only after the last sentence text is already resumable.
plan_append abstractmethod ¶
plan_append(
*,
request_id: str,
fence: DuplexFence,
session_config: dict[str, object],
runtime_config: dict[str, object],
seq: int,
turn_seq: int,
payload: object,
final: bool,
sampling_params: object,
) -> DuplexAppendPlan
plan_partial_stage_output ¶
plan_partial_stage_output(
orchestrator: object,
stage_id: int,
replica_id: int,
output: object,
req_state: object,
) -> PartialStageForward | None
Return a Talker update the orchestrator should submit, or None.
Default models do not split Stage1 text. The orchestrator owns the actual _forward_to_next_stage call.
prepare_append_plan async ¶
prepare_append_plan(**kwargs) -> DuplexAppendPlan
Prepare a plan; plugins may offload expensive work on owned snapshots.
prepare_prompt_config ¶
prepare_prompt_config(
config: dict[str, object],
*,
state: DuplexModelSessionState,
payload: dict[str, object],
) -> dict[str, object]
Add model-owned context before planning an append on the session loop.
prepare_runtime_config abstractmethod async ¶
prepare_runtime_config(
config: DuplexSessionConfig,
*,
model_config: ModelConfig | None,
) -> dict[str, object]
project_intermediate_output ¶
Return True to project this intermediate stage to the client.
Unlike decide_output, projecting does not short-circuit the pipeline: the stage output is still forwarded to the next stage. Default is off. Orthogonal to projects_intermediate_outputs (Qwen3 Stage0); this hook is per-stage.
release_concurrent_turn_requests ¶
release_concurrent_turn_requests(
*,
stage_id: int,
segment_finished: bool,
output: object,
context: object,
) -> bool
Return True when the next user commit may start while prior TTS drains.
The plugin chooses when that is safe. The runner must not hard-code a stage id. Default off. Unlike barge-in, this path must not cancel the old TTS.
runtime_config_for_function_output ¶
runtime_config_for_function_output(
config: DuplexSessionConfig,
current: Mapping[str, object],
item: Mapping[str, object],
) -> dict[str, object] | None
runtime_config_for_update abstractmethod ¶
runtime_config_for_update(
config: DuplexSessionConfig,
current: Mapping[str, object],
) -> dict[str, object]
user_transcript ¶
ASR text to show as the user's words, or None.
Default models do not surface Stage0. AURA uses this for a spoken turn only; vision-follow commits stay off the transcript.
DuplexModelSessionState ¶
DuplexRuntimeConfigError ¶
PartialStageForward dataclass ¶
One downstream update the orchestrator should submit.
close_only is a final update with no new sentence. output is the model-built payload; the orchestrator does not interpret its text.
queue_close_after means this chunk still has text, but Stage1 has finished and an earlier sentence is already in flight. The text must be submitted resumable. A non-resumable submit is an end sentinel (StreamingUpdate.from_request returns None) and aborts that sentence.
PcmAppendBuffer ¶
load_duplex_plugin ¶
load_duplex_plugin(
path: str, encode_audio: EncodeAudio
) -> DuplexModelPlugin
reject_changed_runtime_value ¶
reject_changed_runtime_value(
new_value: object,
current_value: object,
*,
message: str,
code: str,
error_cls: type[
DuplexRuntimeConfigError
] = DuplexRuntimeConfigError,
) -> None
validate_duplex_plugin_sampling ¶
validate_duplex_plugin_sampling(
plugin: DuplexModelPlugin,
*,
sampling_defaults: tuple[object, ...],
) -> None
Fail fast when the plugin cannot produce one sampling parameter per stage.