vllm_omni.worker_v2.omni_model_runner ¶
Omni v2 GPU model runner hooks.
OmniGPUModelRunner ¶
Bases: GPUModelRunner
Thin layer over v2 GPUModelRunner for Omni lifecycle hooks.
capture_model ¶
capture_model() -> int
Handle CUDA graph capture for Omni models.
Tuple-returning models use PIECEWISE graphs; FULL replay requires a tensor output.
For PIECEWISE capture, the warmup pass runs with CUDAGraphMode.NONE which hits torch.empty_like(hidden_states) in the cudagraph framework. If the model returns a tuple, that call crashes. We temporarily wrap the model's forward to extract only the tensor part during capture, then restore the original forward.
execute_model ¶
execute_model(
scheduler_output: SchedulerOutput,
intermediate_tensors: IntermediateTensors | None = None,
dummy_run: bool = False,
skip_attn_for_dummy_run: bool = False,
is_profile: bool = False,
context_len: int = 0,
valid_dummy_state_slots: bool = False,
) -> Any
update_requests ¶
Merge updated additional_information into intermediate_buffer.
In async_chunk mode, chunk_transfer_adapter attaches updated additional_information (e.g. thinker_decode_embeddings) to OmniCachedRequestData for cached requests every schedule step. Upstream GPUModelRunner.update_requests does not handle this field, so we merge it into the intermediate buffer here.