Skip to content

vllm_omni.diffusion.models.minimax_h3.packed_sequence

MiniMax H3 packed-sequence materialization for FL2VA and Ref2VA tasks.

Layout: [text L | imgvid_cond C | audio A(=t*2ch) | video_target V | pad P]. Builder rules: - placeholder ids; block-derived pos infos/masks/cu_seqlens/document_id - img_position_ids fp64 grid: text rows (row_idx,0,0); video/cond t counter continues text_len with temporal interp spans (frame_rescale 5/3 x frame_per_token (1,4,4,4,4)); each spatial sqrt_area axis uses evenly spaced coordinates excluding the right endpoint, then scales them by INTERP; audio channel-major blocks pinned to the w-grid extremes.

MINIMAX_H3_AUDIO_FIRST_ID module-attribute

MINIMAX_H3_AUDIO_FIRST_ID = -15

MINIMAX_H3_AUDIO_ID module-attribute

MINIMAX_H3_AUDIO_ID = -14

MINIMAX_H3_AUDIO_REF_COND_ID module-attribute

MINIMAX_H3_AUDIO_REF_COND_ID = -17

MINIMAX_H3_FL2VA_KEYFRAME_SIGNATURES module-attribute

MINIMAX_H3_FL2VA_KEYFRAME_SIGNATURES: tuple[
    tuple[int, ...], ...
] = ((0,), (-1,), (0, -1))

MINIMAX_H3_IMGVID_COND_ID module-attribute

MINIMAX_H3_IMGVID_COND_ID = -11

MINIMAX_H3_MAX_PAD_SEQ_LEN module-attribute

MINIMAX_H3_MAX_PAD_SEQ_LEN = 1 << 20

Ceiling for a request-pinned packed length.

Every packed row costs a fixed number of bytes across the structural tensors built below, so an unbounded request value would size those allocations. The largest shape this pipeline is documented against -- 672x384 at 209 frames -- packs about 16k rows, so this ceiling leaves roughly two orders of magnitude of headroom while keeping the structural tensors in the tens of megabytes.

MINIMAX_H3_PAD_ID module-attribute

MINIMAX_H3_PAD_ID = -1

MINIMAX_H3_SEQ_ALIGN module-attribute

MINIMAX_H3_SEQ_ALIGN = 64

Row alignment of the packed sequence.

An explicit seq_len must be a multiple of this, so that pinning a length lands on a bucket the default rounding could also produce.

MINIMAX_H3_TEXT_ID module-attribute

MINIMAX_H3_TEXT_ID = -5

MINIMAX_H3_VIDEO_FIRST_ID module-attribute

MINIMAX_H3_VIDEO_FIRST_ID = -3

MINIMAX_H3_VIDEO_ID module-attribute

MINIMAX_H3_VIDEO_ID = -2

MINIMAX_H3_VIDEO_LAST_ID module-attribute

MINIMAX_H3_VIDEO_LAST_ID = -4

minimax_h3_packed_sequence

minimax_h3_packed_sequence(
    *,
    text_len: int,
    latent_t: int,
    latent_h: int,
    latent_w: int,
    audio_t: int,
    audio_channel: int = 2,
    include_keyframe_cond: bool,
    keyframe_frame_indices: list[int]
    | tuple[int, ...]
    | None = None,
    frame_count: int | None = None,
    seq_len: int | None = None,
) -> dict[str, object]

Build the packed-sequence structural fields for one CFG branch.

The used length is padded up to a multiple of :data:MINIMAX_H3_SEQ_ALIGN, or to an explicit seq_len that covers the used rows -- the same contract :func:minimax_h3_packed_sequence_ref2va_blocks already offers. Pinning the length keeps requests of different prompt lengths on one packed shape.

minimax_h3_packed_sequence_ref2va_blocks

minimax_h3_packed_sequence_ref2va_blocks(
    *,
    text_len: int,
    latent_t: int,
    latent_h: int,
    latent_w: int,
    audio_t: int,
    ref_blocks: Sequence[Mapping[str, object]],
    audio_channel: int = 2,
    seq_len: int | None = None,
    temporal_offset: float = 0.0,
    media_time_origin: int | None = None,
) -> dict[str, object]

General ref2va-family packed layout.

ref_blocks are consumed in request/plan order: - {"kind": "image", "latent_h": H, "latent_w": W} - {"kind": "audio", "ref_audio_t": T} - {"kind": "video"|"video_audio", "ref_audio_t": T, "latent_t": RT, "latent_h": RH, "latent_w": RW} - Internal latent_guide: a final AV condition block on the target's spatial grid and temporal origin, without advancing the reference clock.

Video-bearing blocks pack their audio rows immediately before their video rows; both share the same temporal origin and advance by the longer of the audio and video spans. Standalone audio advances the target origin by its own T, and image blocks advance it by one integer slot.

temporal_offset is the window origin in 40-Hz RoPE units. It shifts target AV, reference AV and latent guides together, leaving text, static images, padding and all spatial coordinates unchanged.