vllm_omni.diffusion.models.magi2.preview_data_proxy ¶
Native packing helpers for the MAGI-2 Preview transformer.
The released model consumes one varlen sequence per sample. Tokens are laid out in VIDEO -> AUDIO -> TEXT order, followed by zero or more reference-image special-token/image-token pairs. This module keeps that layout and its 9-D coordinate metadata independent of any particular distributed topology.
The packing math is adapted from SandAI's Apache-2.0 MAGI-2 Preview inference implementation. It intentionally has no dependency on that implementation at runtime.
Magi2DataProxy ¶
Convert dense modality tensors to/from the transformer's packed ABI.
The proxy is request-stateless: all packing state needed by process_output is returned as PackedModelInput.output_layout rather than stored on this shared instance, so concurrent requests cannot overwrite one another.
process_output staticmethod ¶
process_output(
x: Tensor, output_layout: SimplePackedData
) -> tuple[Tensor, Tensor]
Magi2PreviewDataProxyConfig dataclass ¶
Modality ¶
ModelInput dataclass ¶
PackedModelInput dataclass ¶
Request-owned packed transformer arguments and output layout.
SimplePackedData dataclass ¶
A concatenation of independently addressable packed samples.
SingleData dataclass ¶
Packed metadata for one sample.
ref_image_feat_lens class-attribute instance-attribute ¶
ref_image_special_tokens class-attribute instance-attribute ¶
ref_image_special_tokens: list[Tensor] | None = None
spatial_rope_interpolation instance-attribute ¶
spatial_rope_interpolation: Literal['inter', 'extra']
VarlenHandler dataclass ¶
Packed-sequence metadata consumed by MAGI-2 attention.