vllm_omni.model_executor.models.aura_omni.duplex ¶
Modules:
| Name | Description |
|---|---|
capabilities | |
data_plane | Project AURA stage outputs into duplex internal events. |
history | Stage1-local SessionHistory for duplex AURA (import from asr2aura / aura2tts). |
input | Commit-only PCM buffer for AURA duplex (emit whole utterance on commit). |
plugin | AURA full-duplex model plugin: turn-commit four-stage NewRequest. |
sentence_tts | Stage1 sentence handoff onto Talker. AURA-only; the orchestrator just calls it. |
session | |
AuraDuplexPlugin ¶
Bases: DuplexModelPlugin
Turn-commit AURA: one ephemeral four-stage request per utterance.
private_runtime_config_keys class-attribute instance-attribute ¶
commit_model_context ¶
configure_sampling_params ¶
configure_sampling_params(
*,
runtime_config: dict[str, object],
defaults: tuple[object, ...],
) -> tuple[object, ...]
data_plane_context ¶
data_plane_context(
*,
epoch: int,
turn_id: int,
active_response_turn_id: int | None,
active_response_id: str | None,
auto_responds: bool,
response_format: str,
speed: float | None,
modalities: tuple[str, ...],
) -> AuraDataPlaneContext
decide_output ¶
decide_output(
*,
stage_id: int,
final_stage_id: int,
segment_finished: bool,
segment_token_ids: tuple[int, ...],
segment_output_metadata: dict[str, object],
output: object,
) -> DuplexOutputDecision | None
partial_stage_followup ¶
Close the Talker stream after a resumable final sentence.
The sentence text was already forwarded with queue_close_after. Flipping the prompt flag here makes the second submit a close-only sentinel instead of another copy of that sentence.
plan_append ¶
plan_append(
*,
request_id: str,
fence: DuplexFence,
session_config: dict[str, object],
runtime_config: dict[str, object],
seq: int,
turn_seq: int,
payload: object,
final: bool,
sampling_params: object,
) -> DuplexAppendPlan
plan_partial_stage_output ¶
plan_partial_stage_output(
orchestrator: Any,
stage_id: int,
replica_id: int,
output: Any,
req_state: Any,
)
prepare_runtime_config async ¶
prepare_runtime_config(
config: DuplexSessionConfig,
*,
model_config: object | None,
) -> dict[str, object]
project_intermediate_output ¶
Project Stage1 thinker text to the client without short-circuiting TTS.
release_concurrent_turn_requests ¶
release_concurrent_turn_requests(
*,
stage_id: int,
segment_finished: bool,
output: object,
context: object,
) -> bool
After Stage1 text/silent final, next commit may start while TTS drains.
runtime_config_for_update ¶
runtime_config_for_update(
config: DuplexSessionConfig,
current: Mapping[str, object],
) -> dict[str, object]
user_transcript ¶
Stage0 ASR text for a spoken turn, so the demo can show what was said.
Vision-follow sets is_speech false; that audio is a silent pad and must not open a user bubble. The pipeline still forwards Stage0.