Skip to content

vllm_omni.diffusion.models.minimax_h3.batched_packing

Batched multi-request packing for MiniMax H3 step-wise execution.

Request mode forwards exactly one packed sequence per denoise step. Step mode (continuous batching) may hold several requests at once, so this module concatenates their packed layouts into a single sequence and rebuilds cu_seqlens with one document per request plus any nonempty alignment-padding tail. Attention therefore never crosses a request boundary, while every request keeps its own RoPE coordinates, token tags, and timesteps. A batch of one is not rebuilt at all -- it delegates to the request-mode layout.

Row order is the request order, so the concatenated video velocity returned by the DiT can be split back per request by row count alone -- which is exactly what the runner does with StepRequestState.latents.

minimax_h3_batched_forward_kwargs

minimax_h3_batched_forward_kwargs(
    *,
    branches: Sequence[MiniMaxH3DenoiseBranch],
    video_rows: Sequence[Tensor],
    audio_rows: Sequence[Tensor],
    t_video: Sequence[float],
    t_audio: Sequence[float],
    imgvid_cond_timesteps: Sequence[float],
    audio_ref_cond_timesteps: Sequence[float],
    video_target_timesteps: Sequence[Tensor | None]
    | None = None,
    audio_target_timesteps: Sequence[Tensor | None]
    | None = None,
) -> dict[str, Any]

Build one DiT forward kwargs dict covering every request in the batch.

A single request delegates to :meth:MiniMaxH3DenoiseBranch.forward_kwargs, so step mode and request mode are the same call rather than two layouts that happen to agree -- which also keeps video_token_layout (the sparse- attention hint) on the single-request step path.