vllm_omni.engine.mm_outputs ¶
MultimodalCompletionOutput dataclass ¶
Bases: CompletionOutput
CompletionOutput with multimodal support.
Inherits all CompletionOutput fields and adds multimodal_output. As a CompletionOutput subclass, compatible with all existing vLLM consumers.
MultimodalPayload dataclass ¶
Bases: Mapping
Structured multimodal output payload.
Implements collections.abc.Mapping so that isinstance(payload, dict) style checks in downstream code can be replaced with duck-typing, and payload.get(key), payload[key], key in payload, len(payload) all work seamlessly for both tensors and metadata.
Attributes:
| Name | Type | Description |
|---|---|---|
tensors | dict[str, Tensor] | Dictionary mapping modality/key names to their tensors. |
metadata | dict[str, Any] | Optional dictionary for non-tensor metadata (e.g., sample rate for audio, image dimensions). |
metadata class-attribute instance-attribute ¶
primary_tensor property ¶
Return the first tensor in the payload, or None if empty.
tensors class-attribute instance-attribute ¶
consolidate_metadata ¶
Resolve deferred tensor lists in metadata by keeping the latest value.
Metadata values are per-step snapshots (e.g. sample rate), not content deltas, so the latest value supersedes earlier ones. Nested dicts (unflattened payloads) are resolved one level down.
consolidate_tensors ¶
consolidate_tensors(modality: OutputModality) -> None
Concatenate deferred tensor lists into single tensors.
Tensors are generated content accumulated as chunks, so each key's list is concatenated according to the strategy get_accumulation_strategy(modality, key) resolves for it (e.g. audio waveform chunks along the time dimension, latent frames along the batch dimension). Most keys share modality's default, but a key can be registered for a different strategy -- e.g. a codec-frame matrix that grows along dim 0 rather than the waveform-tuned default for its modality -- via register_key_accumulation_strategy.
from_dict classmethod ¶
from_dict(
data: dict[str, Any] | None,
) -> MultimodalPayload | None
Create a MultimodalPayload from a raw dictionary.
Separates torch.Tensor values into tensors and everything else into metadata.
from_raw classmethod ¶
from_raw(
payload: Any, modality_key: str
) -> MultimodalPayload | None
Create a MultimodalPayload from a raw producer payload.
Accepts a MultimodalPayload (returned as-is), a dict, or a bare tensor (stored under modality_key). Tensors are moved to CPU. Producer-specific dict keys are remapped to the semantic modality key (e.g. "audio", "latent"): AR runners produce {"hidden": ...} and generation runners produce {"model_outputs": ...}.
get ¶
Get a value by key, searching tensors first then metadata.
merged_with ¶
merged_with(
incoming: MultimodalPayload,
) -> MultimodalPayload
Merge incoming onto this payload and return the result.
Content tensors accumulate into lists for deferred concatenation; known sample-rate keys are snapshots replaced immediately, including in DELTA streams that do not consolidate each emission. Missing keys retain their previous value. Other values keep the existing merge behavior. When this payload is empty, incoming is returned as-is, so callers should use the return value: accumulated = accumulated.merged_with(incoming).
OutputModality ¶
Bases: Flag
Bit-flag enum for output modalities.
Compose freely with | — no need to enumerate every combination.
Single: OutputModality.TEXT, OutputModality.IMAGE, ... Compound: OutputModality.TEXT | OutputModality.IMAGE (text+image)
Note: POOLING is intentionally excluded. Pooling/embedding is vLLM's native path (pooling_output → PoolingRequestOutput), handled entirely by the base OutputProcessor. vLLM-Omni's layer does not participate.
from_string classmethod ¶
from_string(s: str | None) -> OutputModality
Parse a free-text modality string into an OutputModality flag.
Handles common aliases and compound strings separated by + or ,.
Examples::
OutputModality.from_string("text+image")
# → OutputModality.TEXT | OutputModality.IMAGE
TensorAccumulationStrategy ¶
Bases: Enum
Strategy for merging incremental multimodal tensors.
APPEND_LIST class-attribute instance-attribute ¶
Append to a list (no tensor concatenation).
CONCAT_DIM0 class-attribute instance-attribute ¶
Concatenate along dimension 0. Used for image/latent tensors.
CONCAT_LAST class-attribute instance-attribute ¶
Concatenate along the last dimension. Used for audio waveforms.
REPLACE class-attribute instance-attribute ¶
Replace previous tensor entirely with the latest one.
get_accumulation_strategy ¶
get_accumulation_strategy(
modality: OutputModality, key: str | None = None
) -> TensorAccumulationStrategy
Determine the tensor merge strategy for one output key.
A registered per-key override (see register_key_accumulation_strategy) always wins; otherwise the strategy falls back to the modality-wide default.