Skip to content

vllm_omni.worker_v2.omni_model_runner

Omni v2 GPU model runner hooks.

logger module-attribute

logger = init_logger(__name__)

OmniGPUModelRunner

Bases: GPUModelRunner

Thin layer over v2 GPUModelRunner for Omni lifecycle hooks.

encoder_cache instance-attribute

encoder_cache: EncoderCache | None

model_state instance-attribute

model_state: OmniModelState

supports_mm_inputs instance-attribute

supports_mm_inputs: bool

add_requests

add_requests(scheduler_output: SchedulerOutput) -> None

capture_model

capture_model() -> int

Handle CUDA graph capture for Omni models.

Tuple-returning models use PIECEWISE graphs; FULL replay requires a tensor output.

For PIECEWISE capture, the warmup pass runs with CUDAGraphMode.NONE which hits torch.empty_like(hidden_states) in the cudagraph framework. If the model returns a tuple, that call crashes. We temporarily wrap the model's forward to extract only the tensor part during capture, then restore the original forward.

execute_model

execute_model(
    scheduler_output: SchedulerOutput,
    intermediate_tensors: IntermediateTensors | None = None,
    dummy_run: bool = False,
    skip_attn_for_dummy_run: bool = False,
    is_profile: bool = False,
    context_len: int = 0,
    valid_dummy_state_slots: bool = False,
) -> Any

finish_requests

finish_requests(scheduler_output: SchedulerOutput) -> None

load_model

load_model(*args: Any, **kwargs: Any) -> None

shutdown

shutdown() -> None

update_requests

update_requests(scheduler_output: SchedulerOutput) -> None

Merge updated additional_information into intermediate_buffer.

In async_chunk mode, chunk_transfer_adapter attaches updated additional_information (e.g. thinker_decode_embeddings) to OmniCachedRequestData for cached requests every schedule step. Upstream GPUModelRunner.update_requests does not handle this field, so we merge it into the intermediate buffer here.