vllm_omni.worker_v2.model_states.intermediate_buffer ¶
OmniIntermediateBuffer — per-request cross-stage state for Omni pipelines.
Uses req_index (not req_id) for O(1) access, aligned with v2's RequestState slot management.
OmniIntermediateBuffer ¶
Per-request intermediate state for multi-stage Omni pipelines.
Stores prompt_embeds, additional_information, mm_features, req_id and any runtime updates written back by model postprocess.
gather ¶
Return buffer dicts in current batch order (via idx_mapping_np).
update ¶
update(
req_index: int,
updates: dict[Any, Any],
gpu_resident_keys: set[Any] | None = None,
) -> None
Merge updates into the buffer at req_index.
Tensors are detached; those whose key is not in gpu_resident_keys are moved to CPU.
update_gpu_tensor_rows ¶
update_gpu_tensor_rows(
req_indices: list[int],
key: Any,
values: Tensor,
*,
keepdim: bool = True,
) -> None
Snapshot a batch-first tensor once and retain owned row views.
update_owned_gpu_tensor_rows ¶
update_owned_gpu_tensor_rows(
req_indices: list[int],
key: Any,
owned_values: Tensor,
*,
keepdim: bool = True,
) -> None
Store row views of an already-owned batch tensor without cloning.
The caller must guarantee owned_values is not written after this call (the model output ownership contract). Row views keep the owned storage alive until each request row is replaced or the slot is freed.