vllm_omni.model_executor.models.minimax_h3.preprocessing ¶
Shared MiniMax H3 media normalization and Qwen presentation building.
Builds the positive presentation token stream: - fl2va: '
All presentation variants are emitted through the shared _Presentation accumulator so ids and AdaLN token tags cannot drift apart.
MINIMAX_H3_MAX_REFERENCE_IMAGE_BYTES module-attribute ¶
MINIMAX_H3_REFERENCE_IMAGE_FORMATS module-attribute ¶
MINIMAX_H3_REFERENCE_IMAGE_FORMATS = frozenset(
{"jpeg", "png", "webp", "heic", "heif"}
)
MINIMAX_H3_REFERENCE_IMAGE_SHORT_EDGE module-attribute ¶
MINIMAX_H3_SUPPORTED_ASPECT_RATIOS module-attribute ¶
MINIMAX_H3_SUPPORTED_ASPECT_RATIOS = {
"21:9": 21.0 / 9.0,
"16:9": 16.0 / 9.0,
"4:3": 4.0 / 3.0,
"1:1": 1.0,
"3:4": 3.0 / 4.0,
"9:16": 9.0 / 16.0,
}
build_minimax_h3_presentation ¶
build_minimax_h3_presentation(
tokenizer: Any,
*,
prompt: str,
task: str,
condition_labels: list[tuple[str, int]],
image_grid_thw: Tensor | None,
video_grid_thw: Tensor | None,
video_timestamps: Sequence[Sequence[float]] | None,
merge_size: int,
) -> tuple[Tensor, Tensor]
Build the token IDs and role IDs shared by fused and split H3.
load_minimax_h3_images ¶
Normalize one or more H3 image inputs to RGB PIL images.
minimax_h3_multi_image_presentation ¶
minimax_h3_multi_image_presentation(
tokenizer: Any,
*,
prompt: str,
image_token_counts: list[int],
) -> tuple[Tensor, Tensor]
Return aligned FL2VA presentation token IDs and role IDs.
minimax_h3_ref2va_presentation ¶
minimax_h3_ref2va_presentation(
tokenizer: Any,
*,
prompt: str,
condition_labels: list[tuple[str, int]],
image_token_count: int | list[int] | None,
) -> tuple[Tensor, Tensor]
ref2va positive presentation:
per condition in request order — image i: <Picture i>: label followed by the vision block; audio j: <Audio j>: label only (audio content never enters Qwen) — then the verbatim prompt. Returns (ids, token_tags) with the vision block tagged VIDEO(0) and everything else TEXT(1).
condition_labels: [("image", 1), ("audio", 1), ...] with 1-based ordinals per type.
minimax_h3_ref2va_video_presentation ¶
minimax_h3_ref2va_video_presentation(
tokenizer: Any,
*,
prompt: str,
condition_labels: list[tuple[str, int]],
image_token_count: int | list[int] | None,
video_block_token_counts: list[int]
| list[list[int]]
| None,
video_block_timestamps: list[float]
| list[list[float]]
| None,
) -> tuple[Tensor, Tensor]
ref2va (optionally with video refs) positive presentation:
per condition in request order — - image i: <Picture i>: label + one image vision block; - audio j: <Audio j>: label only (audio content never enters Qwen); - video k: <Video k>: label, then per temporal block a timestamp text <{t:.1f} seconds> followed by a VIDEO vision block (<|vision_start|> + <|video_pad|> x n + <|vision_end|>). Timestamps are the mean of each merged frame pair (Qwen3VL temporal merge 2; odd frame counts repeat the last frame), emitting the <0.2 seconds> .. <4.0 seconds> sequence — note Python bankers-rounding at .1f. then the verbatim prompt. Vision blocks are tagged VIDEO(0), everything else TEXT(1).
minimax_h3_text_only_ids ¶
t2va presentation: verbatim prompt, no special tokens.
resolve_minimax_h3_aspect_ratio ¶
Resolve H3's task-specific output aspect-ratio policy.
resolve_minimax_h3_output_canvas ¶
Resolve the official H3 ratio/area policy to a 32-pixel canvas.