Skip to content

vllm_omni.model_executor.models.minimax_h3.preprocessing

Shared MiniMax H3 media normalization and Qwen presentation building.

Builds the positive presentation token stream: - fl2va: ': ' label + vision block (<|vision_start|> + N*<|image_pad|> + <|vision_end|>) + prompt text. - t2va: prompt text only (no vision block). Prompt text passes through verbatim (no stripping or rewriting).

All presentation variants are emitted through the shared _Presentation accumulator so ids and AdaLN token tags cannot drift apart.

IMAGE_PAD module-attribute

IMAGE_PAD = '<|image_pad|>'

MINIMAX_H3_MAX_REFERENCE_IMAGE_BYTES module-attribute

MINIMAX_H3_MAX_REFERENCE_IMAGE_BYTES = 30 * 1024 * 1024

MINIMAX_H3_OUTPUT_MAX_PIXELS module-attribute

MINIMAX_H3_OUTPUT_MAX_PIXELS = 768 * 1344

MINIMAX_H3_OUTPUT_SHORT_EDGE module-attribute

MINIMAX_H3_OUTPUT_SHORT_EDGE = 768

MINIMAX_H3_REFERENCE_IMAGE_FORMATS module-attribute

MINIMAX_H3_REFERENCE_IMAGE_FORMATS = frozenset(
    {"jpeg", "png", "webp", "heic", "heif"}
)

MINIMAX_H3_REFERENCE_IMAGE_MULTIPLE module-attribute

MINIMAX_H3_REFERENCE_IMAGE_MULTIPLE = 32

MINIMAX_H3_REFERENCE_IMAGE_SHORT_EDGE module-attribute

MINIMAX_H3_REFERENCE_IMAGE_SHORT_EDGE = 2048

MINIMAX_H3_SUPPORTED_ASPECT_RATIOS module-attribute

MINIMAX_H3_SUPPORTED_ASPECT_RATIOS = {
    "21:9": 21.0 / 9.0,
    "16:9": 16.0 / 9.0,
    "4:3": 4.0 / 3.0,
    "1:1": 1.0,
    "3:4": 3.0 / 4.0,
    "9:16": 9.0 / 16.0,
}

VIDEO_PAD module-attribute

VIDEO_PAD = '<|video_pad|>'

VISION_END module-attribute

VISION_END = '<|vision_end|>'

VISION_START module-attribute

VISION_START = '<|vision_start|>'

build_minimax_h3_presentation

build_minimax_h3_presentation(
    tokenizer: Any,
    *,
    prompt: str,
    task: str,
    condition_labels: list[tuple[str, int]],
    image_grid_thw: Tensor | None,
    video_grid_thw: Tensor | None,
    video_timestamps: Sequence[Sequence[float]] | None,
    merge_size: int,
) -> tuple[Tensor, Tensor]

Build the token IDs and role IDs shared by fused and split H3.

load_minimax_h3_images

load_minimax_h3_images(value: Any) -> list[Image]

Normalize one or more H3 image inputs to RGB PIL images.

minimax_h3_multi_image_presentation

minimax_h3_multi_image_presentation(
    tokenizer: Any,
    *,
    prompt: str,
    image_token_counts: list[int],
) -> tuple[Tensor, Tensor]

Return aligned FL2VA presentation token IDs and role IDs.

minimax_h3_ref2va_presentation

minimax_h3_ref2va_presentation(
    tokenizer: Any,
    *,
    prompt: str,
    condition_labels: list[tuple[str, int]],
    image_token_count: int | list[int] | None,
) -> tuple[Tensor, Tensor]

ref2va positive presentation:

per condition in request order — image i: <Picture i>: label followed by the vision block; audio j: <Audio j>: label only (audio content never enters Qwen) — then the verbatim prompt. Returns (ids, token_tags) with the vision block tagged VIDEO(0) and everything else TEXT(1).

condition_labels: [("image", 1), ("audio", 1), ...] with 1-based ordinals per type.

minimax_h3_ref2va_video_presentation

minimax_h3_ref2va_video_presentation(
    tokenizer: Any,
    *,
    prompt: str,
    condition_labels: list[tuple[str, int]],
    image_token_count: int | list[int] | None,
    video_block_token_counts: list[int]
    | list[list[int]]
    | None,
    video_block_timestamps: list[float]
    | list[list[float]]
    | None,
) -> tuple[Tensor, Tensor]

ref2va (optionally with video refs) positive presentation:

per condition in request order — - image i: <Picture i>: label + one image vision block; - audio j: <Audio j>: label only (audio content never enters Qwen); - video k: <Video k>: label, then per temporal block a timestamp text <{t:.1f} seconds> followed by a VIDEO vision block (<|vision_start|> + <|video_pad|> x n + <|vision_end|>). Timestamps are the mean of each merged frame pair (Qwen3VL temporal merge 2; odd frame counts repeat the last frame), emitting the <0.2 seconds> .. <4.0 seconds> sequence — note Python bankers-rounding at .1f. then the verbatim prompt. Vision blocks are tagged VIDEO(0), everything else TEXT(1).

minimax_h3_text_only_ids

minimax_h3_text_only_ids(
    tokenizer: Any, prompt: str
) -> Tensor

t2va presentation: verbatim prompt, no special tokens.

resolve_minimax_h3_aspect_ratio

resolve_minimax_h3_aspect_ratio(
    task: str, value: Any, image: Image | None
) -> float

Resolve H3's task-specific output aspect-ratio policy.

resolve_minimax_h3_output_canvas

resolve_minimax_h3_output_canvas(
    aspect_ratio: float, short_edge: int
) -> tuple[int, int]

Resolve the official H3 ratio/area policy to a 32-pixel canvas.

resolve_minimax_h3_reference_image_shape

resolve_minimax_h3_reference_image_shape(
    image: Image,
) -> tuple[int, int]

Resize an H3 reference image to the official 2048-short-edge canvas.