vllm_omni.worker.output.payload_build ¶
Per-request multimodal payload builders.
State-free extraction helpers that turn a step's multimodal_outputs (combined prefix-cache-merged form, or the direct mm_cpu form) into one request's payload dict. Sparse-routing context (audio_sparse_output, sparse_mm_index) is resolved upstream by vllm_omni.worker.sparse_audio.resolve_sparse_mm_routing and arrives here as plain arguments.
build_combined_prefix_cache_mm_payload ¶
build_combined_prefix_cache_mm_payload(
combined_multimodal_outputs: dict,
*,
rid: str,
list_idx: int,
) -> dict[str, object]
Unwrap one request's entry from the prefix-cache-merged payload.
Sparse-agnostic by design: list_idx arrives pre-resolved by the caller (build_omni_mm_payload, which owns the sparse-routing context) — batch index normally, sparse index under sparse routing.
build_omni_mm_payload ¶
build_omni_mm_payload(
*,
combined_multimodal_outputs: dict | None,
mm_cpu: dict[str, object] | None,
rid: str,
idx: int,
start: int,
end: int,
audio_sparse_output: bool,
sparse_mm_index: dict[str, int],
hidden_seq_len: int,
scheduled_seq_len: int,
) -> dict[str, object]
Build one request's multimodal payload for the step.
unwrap_combined_payload_value ¶
unwrap_combined_payload_value(
value: Any, *, rid: str, list_idx: int, mm_key: str
) -> tuple[bool, Any]
Unwrap one value from the prefix-cache-merged payload.
Returns (keep, unwrapped): keep is False when a non-singleton per-request list is misaligned and must be dropped. A singleton list is request-invariant passthrough data and is shared across the batch.