vllm_omni.diffusion.models.minimax_h3.packed_sequence ¶
MiniMax H3 packed-sequence materialization for FL2VA and Ref2VA tasks.
Layout: [text L | imgvid_cond C | audio A(=t*2ch) | video_target V | pad P]. Builder rules: - placeholder ids; block-derived pos infos/masks/cu_seqlens/document_id - img_position_ids fp64 grid: text rows (row_idx,0,0); video/cond t counter continues text_len with temporal interp spans (frame_rescale 5/3 x frame_per_token (1,4,4,4,4)); each spatial sqrt_area axis uses evenly spaced coordinates excluding the right endpoint, then scales them by INTERP; audio channel-major blocks pinned to the w-grid extremes.
MINIMAX_H3_FL2VA_KEYFRAME_SIGNATURES module-attribute ¶
MINIMAX_H3_MAX_PAD_SEQ_LEN module-attribute ¶
Ceiling for a request-pinned packed length.
Every packed row costs a fixed number of bytes across the structural tensors built below, so an unbounded request value would size those allocations. The largest shape this pipeline is documented against -- 672x384 at 209 frames -- packs about 16k rows, so this ceiling leaves roughly two orders of magnitude of headroom while keeping the structural tensors in the tens of megabytes.
MINIMAX_H3_SEQ_ALIGN module-attribute ¶
Row alignment of the packed sequence.
An explicit seq_len must be a multiple of this, so that pinning a length lands on a bucket the default rounding could also produce.
minimax_h3_packed_sequence ¶
minimax_h3_packed_sequence(
*,
text_len: int,
latent_t: int,
latent_h: int,
latent_w: int,
audio_t: int,
audio_channel: int = 2,
include_keyframe_cond: bool,
keyframe_frame_indices: list[int]
| tuple[int, ...]
| None = None,
frame_count: int | None = None,
seq_len: int | None = None,
) -> dict[str, object]
Build the packed-sequence structural fields for one CFG branch.
The used length is padded up to a multiple of :data:MINIMAX_H3_SEQ_ALIGN, or to an explicit seq_len that covers the used rows -- the same contract :func:minimax_h3_packed_sequence_ref2va_blocks already offers. Pinning the length keeps requests of different prompt lengths on one packed shape.
minimax_h3_packed_sequence_ref2va_blocks ¶
minimax_h3_packed_sequence_ref2va_blocks(
*,
text_len: int,
latent_t: int,
latent_h: int,
latent_w: int,
audio_t: int,
ref_blocks: Sequence[Mapping[str, object]],
audio_channel: int = 2,
seq_len: int | None = None,
temporal_offset: float = 0.0,
media_time_origin: int | None = None,
) -> dict[str, object]
General ref2va-family packed layout.
ref_blocks are consumed in request/plan order: - {"kind": "image", "latent_h": H, "latent_w": W} - {"kind": "audio", "ref_audio_t": T} - {"kind": "video"|"video_audio", "ref_audio_t": T, "latent_t": RT, "latent_h": RH, "latent_w": RW} - Internal latent_guide: a final AV condition block on the target's spatial grid and temporal origin, without advancing the reference clock.
Video-bearing blocks pack their audio rows immediately before their video rows; both share the same temporal origin and advance by the longer of the audio and video spans. Standalone audio advances the target origin by its own T, and image blocks advance it by one integer slot.
temporal_offset is the window origin in 40-Hz RoPE units. It shifts target AV, reference AV and latent guides together, leaving text, static images, padding and all spatial coordinates unchanged.