vllm_omni.worker_v2.omni_generation_model_runner ¶
OmniGenerationModelRunner — non-autoregressive stage runner on MR V2.
Used for stages like Code2Wav that convert codec codes to audio waveforms. No token sampling or logits computation — model output goes directly into pooler_output. Inherits from OmniGPUModelRunner for intermediate buffer and lifecycle hooks.
OmniGenerationAsyncOutput ¶
Bases: AsyncModelRunnerOutput
Overlap generation-stage D2H copies with the next model step.
multimodal_outputs_cpu instance-attribute ¶
multimodal_outputs_cpu = _async_copy_mm(
multimodal_outputs,
total_tokens=0,
copy_stream=copy_stream,
pin_memory=PIN_MEMORY,
)
OmniGenerationModelRunner ¶
Bases: OmniGPUModelRunner
Non-autoregressive generation runner (e.g. Code2Wav).
Overrides execute_model to skip the tensor-only assertion and sample_tokens to construct pooler_output from multimodal model outputs without performing token sampling.
execute_model ¶
execute_model(
scheduler_output: SchedulerOutput,
intermediate_tensors: IntermediateTensors | None = None,
dummy_run: bool = False,
skip_attn_for_dummy_run: bool = False,
is_profile: bool = False,
context_len: int = 0,
valid_dummy_state_slots: bool = False,
) -> ModelRunnerOutput | IntermediateTensors | None
profile_run ¶
Generation models have no KV cache — skip profiling.
Code2Wav shares GPU memory with the Talker stage (same device); its memory footprint is managed via gpu_memory_utilization config, not profiled dynamically. Running the real model with random input_ids causes out-of-bounds indexing in codec lookup tables.
sample_tokens ¶
sample_tokens(
grammar_output: GrammarOutput | None = None,
) -> OmniModelRunnerOutput | AsyncModelRunnerOutput | None