Skip to content

vllm_omni.worker_v2.model_states.intermediate_buffer

OmniIntermediateBuffer — per-request cross-stage state for Omni pipelines.

Uses req_index (not req_id) for O(1) access, aligned with v2's RequestState slot management.

OmniIntermediateBuffer

Per-request intermediate state for multi-stage Omni pipelines.

Stores prompt_embeds, additional_information, mm_features, req_id and any runtime updates written back by model postprocess.

buffers instance-attribute

buffers: list[dict[str, Any]] = [
    {} for _ in range(max_num_reqs)
]

req_id_to_index instance-attribute

req_id_to_index: dict[str, int] = {}

add_request

add_request(
    req_index: int, new_req_data: NewRequestData
) -> None

gather

gather(input_batch: InputBatch) -> list[dict[str, Any]]

Return buffer dicts in current batch order (via idx_mapping_np).

remove_request

remove_request(req_index: int) -> None

update

update(
    req_index: int,
    updates: dict[Any, Any],
    gpu_resident_keys: set[Any] | None = None,
) -> None

Merge updates into the buffer at req_index.

Tensors are detached; those whose key is not in gpu_resident_keys are moved to CPU.

update_gpu_tensor_rows

update_gpu_tensor_rows(
    req_indices: list[int],
    key: Any,
    values: Tensor,
    *,
    keepdim: bool = True,
) -> None

Snapshot a batch-first tensor once and retain owned row views.

update_owned_gpu_tensor_rows

update_owned_gpu_tensor_rows(
    req_indices: list[int],
    key: Any,
    owned_values: Tensor,
    *,
    keepdim: bool = True,
) -> None

Store row views of an already-owned batch tensor without cloning.

The caller must guarantee owned_values is not written after this call (the model output ownership contract). Row views keep the owned storage alive until each request row is replaced or the slot is freed.