Skip to content

vllm_omni.worker_v2.omni_generation_model_runner

OmniGenerationModelRunner — non-autoregressive stage runner on MR V2.

Used for stages like Code2Wav that convert codec codes to audio waveforms. No token sampling or logits computation — model output goes directly into pooler_output. Inherits from OmniGPUModelRunner for intermediate buffer and lifecycle hooks.

logger module-attribute

logger = init_logger(__name__)

OmniGenerationAsyncOutput

Bases: AsyncModelRunnerOutput

Overlap generation-stage D2H copies with the next model step.

copy_event instance-attribute

copy_event = torch.cuda.Event(blocking=True)

finalize_output instance-attribute

finalize_output = finalize_output

model_runner_output instance-attribute

model_runner_output = model_runner_output

multimodal_outputs_cpu instance-attribute

multimodal_outputs_cpu = _async_copy_mm(
    multimodal_outputs,
    total_tokens=0,
    copy_stream=copy_stream,
    pin_memory=PIN_MEMORY,
)

num_reqs instance-attribute

num_reqs = num_reqs

get_output

get_output() -> OmniModelRunnerOutput

OmniGenerationModelRunner

Bases: OmniGPUModelRunner

Non-autoregressive generation runner (e.g. Code2Wav).

Overrides execute_model to skip the tensor-only assertion and sample_tokens to construct pooler_output from multimodal model outputs without performing token sampling.

execute_model

execute_model(
    scheduler_output: SchedulerOutput,
    intermediate_tensors: IntermediateTensors | None = None,
    dummy_run: bool = False,
    skip_attn_for_dummy_run: bool = False,
    is_profile: bool = False,
    context_len: int = 0,
    valid_dummy_state_slots: bool = False,
) -> ModelRunnerOutput | IntermediateTensors | None

profile_run

profile_run() -> None

Generation models have no KV cache — skip profiling.

Code2Wav shares GPU memory with the Talker stage (same device); its memory footprint is managed via gpu_memory_utilization config, not profiled dynamically. Running the real model with random input_ids causes out-of-bounds indexing in codec lookup tables.

sample_tokens

sample_tokens(
    grammar_output: GrammarOutput | None = None,
) -> OmniModelRunnerOutput | AsyncModelRunnerOutput | None