Skip to content

vllm_omni.engine.mm_outputs

logger module-attribute

logger = init_logger(__name__)

MultimodalCompletionOutput dataclass

Bases: CompletionOutput

CompletionOutput with multimodal support.

Inherits all CompletionOutput fields and adds multimodal_output. As a CompletionOutput subclass, compatible with all existing vLLM consumers.

multimodal_output instance-attribute

multimodal_output = multimodal_output

MultimodalPayload dataclass

Bases: Mapping

Structured multimodal output payload.

Implements collections.abc.Mapping so that isinstance(payload, dict) style checks in downstream code can be replaced with duck-typing, and payload.get(key), payload[key], key in payload, len(payload) all work seamlessly for both tensors and metadata.

Attributes:

Name Type Description
tensors dict[str, Tensor]

Dictionary mapping modality/key names to their tensors.

metadata dict[str, Any]

Optional dictionary for non-tensor metadata (e.g., sample rate for audio, image dimensions).

is_empty property

is_empty: bool

Return True if the payload has no tensors and no metadata.

metadata class-attribute instance-attribute

metadata: dict[str, Any] = field(default_factory=dict)

primary_tensor property

primary_tensor: Tensor | None

Return the first tensor in the payload, or None if empty.

tensors class-attribute instance-attribute

tensors: dict[str, Tensor] = field(default_factory=dict)

consolidate_metadata

consolidate_metadata() -> None

Resolve deferred tensor lists in metadata by keeping the latest value.

Metadata values are per-step snapshots (e.g. sample rate), not content deltas, so the latest value supersedes earlier ones. Nested dicts (unflattened payloads) are resolved one level down.

consolidate_tensors

consolidate_tensors(modality: OutputModality) -> None

Concatenate deferred tensor lists into single tensors.

Tensors are generated content accumulated as chunks, so each key's list is concatenated according to the strategy get_accumulation_strategy(modality, key) resolves for it (e.g. audio waveform chunks along the time dimension, latent frames along the batch dimension). Most keys share modality's default, but a key can be registered for a different strategy -- e.g. a codec-frame matrix that grows along dim 0 rather than the waveform-tuned default for its modality -- via register_key_accumulation_strategy.

from_dict classmethod

from_dict(
    data: dict[str, Any] | None,
) -> MultimodalPayload | None

Create a MultimodalPayload from a raw dictionary.

Separates torch.Tensor values into tensors and everything else into metadata.

from_raw classmethod

from_raw(
    payload: Any, modality_key: str
) -> MultimodalPayload | None

Create a MultimodalPayload from a raw producer payload.

Accepts a MultimodalPayload (returned as-is), a dict, or a bare tensor (stored under modality_key). Tensors are moved to CPU. Producer-specific dict keys are remapped to the semantic modality key (e.g. "audio", "latent"): AR runners produce {"hidden": ...} and generation runners produce {"model_outputs": ...}.

get

get(key: str, default: Any = None) -> Any

Get a value by key, searching tensors first then metadata.

merged_with

merged_with(
    incoming: MultimodalPayload,
) -> MultimodalPayload

Merge incoming onto this payload and return the result.

Content tensors accumulate into lists for deferred concatenation; known sample-rate keys are snapshots replaced immediately, including in DELTA streams that do not consolidate each emission. Missing keys retain their previous value. Other values keep the existing merge behavior. When this payload is empty, incoming is returned as-is, so callers should use the return value: accumulated = accumulated.merged_with(incoming).

to_dict

to_dict() -> dict[str, Any]

Convert back to a plain dict (tensors + metadata merged).

OutputModality

Bases: Flag

Bit-flag enum for output modalities.

Compose freely with | — no need to enumerate every combination.

Single: OutputModality.TEXT, OutputModality.IMAGE, ... Compound: OutputModality.TEXT | OutputModality.IMAGE (text+image)

Note: POOLING is intentionally excluded. Pooling/embedding is vLLM's native path (pooling_output → PoolingRequestOutput), handled entirely by the base OutputProcessor. vLLM-Omni's layer does not participate.

AUDIO class-attribute instance-attribute

AUDIO = auto()

IMAGE class-attribute instance-attribute

IMAGE = auto()

LATENT class-attribute instance-attribute

LATENT = auto()

TEXT class-attribute instance-attribute

TEXT = auto()

has_multimodal property

has_multimodal: bool

has_text property

has_text: bool

from_string classmethod

from_string(s: str | None) -> OutputModality

Parse a free-text modality string into an OutputModality flag.

Handles common aliases and compound strings separated by + or ,.

Examples::

OutputModality.from_string("text+image")
# → OutputModality.TEXT | OutputModality.IMAGE

TensorAccumulationStrategy

Bases: Enum

Strategy for merging incremental multimodal tensors.

APPEND_LIST class-attribute instance-attribute

APPEND_LIST = 'append_list'

Append to a list (no tensor concatenation).

CONCAT_DIM0 class-attribute instance-attribute

CONCAT_DIM0 = 'concat_dim0'

Concatenate along dimension 0. Used for image/latent tensors.

CONCAT_LAST class-attribute instance-attribute

CONCAT_LAST = 'concat_last'

Concatenate along the last dimension. Used for audio waveforms.

REPLACE class-attribute instance-attribute

REPLACE = 'replace'

Replace previous tensor entirely with the latest one.

get_accumulation_strategy

get_accumulation_strategy(
    modality: OutputModality, key: str | None = None
) -> TensorAccumulationStrategy

Determine the tensor merge strategy for one output key.

A registered per-key override (see register_key_accumulation_strategy) always wins; otherwise the strategy falls back to the modality-wide default.