vllm_omni.worker.output ¶
Worker output helpers.
Package seeded ahead of the async-output refactor (RFC #5450 series, G5-C2), which relocates the remaining payload-copy helpers, snapshot types, and ExecuteModelState here. Currently hosts the per-request multimodal payload builders.
Modules:
| Name | Description |
|---|---|
payload_build | Per-request multimodal payload builders. |
build_combined_prefix_cache_mm_payload ¶
build_combined_prefix_cache_mm_payload(
combined_multimodal_outputs: dict,
*,
rid: str,
list_idx: int,
) -> dict[str, object]
Unwrap one request's entry from the prefix-cache-merged payload.
Sparse-agnostic by design: list_idx arrives pre-resolved by the caller (build_omni_mm_payload, which owns the sparse-routing context) — batch index normally, sparse index under sparse routing.
build_omni_mm_payload ¶
build_omni_mm_payload(
*,
combined_multimodal_outputs: dict | None,
mm_cpu: dict[str, object] | None,
rid: str,
idx: int,
start: int,
end: int,
audio_sparse_output: bool,
sparse_mm_index: dict[str, int],
hidden_seq_len: int,
scheduled_seq_len: int,
) -> dict[str, object]
Build one request's multimodal payload for the step.
unwrap_combined_payload_value ¶
unwrap_combined_payload_value(
value: Any, *, rid: str, list_idx: int, mm_key: str
) -> tuple[bool, Any]
Unwrap one value from the prefix-cache-merged payload.
Returns (keep, unwrapped): keep is False when a non-singleton per-request list is misaligned and must be dropped. A singleton list is request-invariant passthrough data and is shared across the batch.