Skip to content

vllm_omni.model_executor.duplex_sampling

DuplexSamplingHelper

Tracks duplex sampling rows outside the generic AR model runner.

active_request_ids instance-attribute

active_request_ids: set[str] = set()

hook_active instance-attribute

hook_active = False

clear

clear() -> None

refresh_active_request

refresh_active_request(runner: object, req_id: str) -> None

rows

rows(runner: object) -> tuple[DuplexSamplingRow, ...]

update_states

update_states(
    runner: object, scheduler_output: object
) -> None

DuplexSamplingRow dataclass

Request-local context for the experimental duplex sampling hook.

max_tokens instance-attribute

max_tokens: int | None

payload instance-attribute

payload: dict[str, object] | None

request_id instance-attribute

request_id: str

row_idx instance-attribute

row_idx: int

seq instance-attribute

seq: int | None

session_id instance-attribute

session_id: str | None

DuplexSamplingRunnerMixin

Runner-side wiring for the experimental duplex sampling hook.

GPU and NPU AR runners are siblings, not parent and child, so neither inherits the other's _sample. Keep the hook here and have each runner opt in, instead of copying the GPU _sample path and dropping the prepare_duplex_sampling call.