Skip to content

vllm_omni.worker.output.payload_build

Per-request multimodal payload builders.

State-free extraction helpers that turn a step's multimodal_outputs (combined prefix-cache-merged form, or the direct mm_cpu form) into one request's payload dict. Sparse-routing context (audio_sparse_output, sparse_mm_index) is resolved upstream by vllm_omni.worker.sparse_audio.resolve_sparse_mm_routing and arrives here as plain arguments.

logger module-attribute

logger = init_logger(__name__)

build_combined_prefix_cache_mm_payload

build_combined_prefix_cache_mm_payload(
    combined_multimodal_outputs: dict,
    *,
    rid: str,
    list_idx: int,
) -> dict[str, object]

Unwrap one request's entry from the prefix-cache-merged payload.

Sparse-agnostic by design: list_idx arrives pre-resolved by the caller (build_omni_mm_payload, which owns the sparse-routing context) — batch index normally, sparse index under sparse routing.

build_omni_mm_payload

build_omni_mm_payload(
    *,
    combined_multimodal_outputs: dict | None,
    mm_cpu: dict[str, object] | None,
    rid: str,
    idx: int,
    start: int,
    end: int,
    audio_sparse_output: bool,
    sparse_mm_index: dict[str, int],
    hidden_seq_len: int,
    scheduled_seq_len: int,
) -> dict[str, object]

Build one request's multimodal payload for the step.

unwrap_combined_payload_value

unwrap_combined_payload_value(
    value: Any, *, rid: str, list_idx: int, mm_key: str
) -> tuple[bool, Any]

Unwrap one value from the prefix-cache-merged payload.

Returns (keep, unwrapped): keep is False when a non-singleton per-request list is misaligned and must be dropped. A singleton list is request-invariant passthrough data and is shared across the batch.