Skip to content

vllm_omni.model_executor.models.aura_omni.duplex

Modules:

Name Description
capabilities
data_plane

Project AURA stage outputs into duplex internal events.

history

Stage1-local SessionHistory for duplex AURA (import from asr2aura / aura2tts).

input

Commit-only PCM buffer for AURA duplex (emit whole utterance on commit).

plugin

AURA full-duplex model plugin: turn-commit four-stage NewRequest.

sentence_tts

Stage1 sentence handoff onto Talker. AURA-only; the orchestrator just calls it.

session

AuraDuplexPlugin

Bases: DuplexModelPlugin

Turn-commit AURA: one ephemeral four-stage request per utterance.

data_plane instance-attribute

data_plane = AuraDataPlaneSession(encode_audio)

plugin_id class-attribute instance-attribute

plugin_id = 'aura'

private_runtime_config_keys class-attribute instance-attribute

private_runtime_config_keys = _PRIVATE_KEYS

capabilities

capabilities(*, max_sessions: int) -> DuplexCapabilities

commit_model_context

commit_model_context(
    *, session_id: str | None, assistant_text: str
) -> None

configure_sampling_params

configure_sampling_params(
    *,
    runtime_config: dict[str, object],
    defaults: tuple[object, ...],
) -> tuple[object, ...]

create_session_state

create_session_state() -> AuraServingSessionState

data_plane_context

data_plane_context(
    *,
    epoch: int,
    turn_id: int,
    active_response_turn_id: int | None,
    active_response_id: str | None,
    auto_responds: bool,
    response_format: str,
    speed: float | None,
    modalities: tuple[str, ...],
) -> AuraDataPlaneContext

decide_output

decide_output(
    *,
    stage_id: int,
    final_stage_id: int,
    segment_finished: bool,
    segment_token_ids: tuple[int, ...],
    segment_output_metadata: dict[str, object],
    output: object,
) -> DuplexOutputDecision | None

draining_stage_ids

draining_stage_ids(*, stage_count: int) -> frozenset[int]

partial_stage_followup

partial_stage_followup(plan: Any, req_state: Any) -> Any

Close the Talker stream after a resumable final sentence.

The sentence text was already forwarded with queue_close_after. Flipping the prompt flag here makes the second submit a close-only sentinel instead of another copy of that sentence.

plan_append

plan_append(
    *,
    request_id: str,
    fence: DuplexFence,
    session_config: dict[str, object],
    runtime_config: dict[str, object],
    seq: int,
    turn_seq: int,
    payload: object,
    final: bool,
    sampling_params: object,
) -> DuplexAppendPlan

plan_partial_stage_output

plan_partial_stage_output(
    orchestrator: Any,
    stage_id: int,
    replica_id: int,
    output: Any,
    req_state: Any,
)

prepare_runtime_config async

prepare_runtime_config(
    config: DuplexSessionConfig,
    *,
    model_config: object | None,
) -> dict[str, object]

project_intermediate_output

project_intermediate_output(
    *, stage_id: int, output: object, context: object
) -> bool

Project Stage1 thinker text to the client without short-circuiting TTS.

release_concurrent_turn_requests

release_concurrent_turn_requests(
    *,
    stage_id: int,
    segment_finished: bool,
    output: object,
    context: object,
) -> bool

After Stage1 text/silent final, next commit may start while TTS drains.

runtime_config_for_update

runtime_config_for_update(
    config: DuplexSessionConfig,
    current: Mapping[str, object],
) -> dict[str, object]

user_transcript

user_transcript(
    *,
    stage_id: int,
    output: object,
    prompt: object,
    finished: bool,
) -> str | None

Stage0 ASR text for a spoken turn, so the demo can show what was said.

Vision-follow sets is_speech false; that audio is a silent pad and must not open a user bubble. The pipeline still forwards Stage0.

validate_client_extra_body

validate_client_extra_body(extra_body: object) -> None