vllm_omni.diffusion.models.minimax_h3.batched_packing ¶
Batched multi-request packing for MiniMax H3 step-wise execution.
Request mode forwards exactly one packed sequence per denoise step. Step mode (continuous batching) may hold several requests at once, so this module concatenates their packed layouts into a single sequence and rebuilds cu_seqlens with one document per request plus any nonempty alignment-padding tail. Attention therefore never crosses a request boundary, while every request keeps its own RoPE coordinates, token tags, and timesteps. A batch of one is not rebuilt at all -- it delegates to the request-mode layout.
Row order is the request order, so the concatenated video velocity returned by the DiT can be split back per request by row count alone -- which is exactly what the runner does with StepRequestState.latents.
minimax_h3_batched_forward_kwargs ¶
minimax_h3_batched_forward_kwargs(
*,
branches: Sequence[MiniMaxH3DenoiseBranch],
video_rows: Sequence[Tensor],
audio_rows: Sequence[Tensor],
t_video: Sequence[float],
t_audio: Sequence[float],
imgvid_cond_timesteps: Sequence[float],
audio_ref_cond_timesteps: Sequence[float],
video_target_timesteps: Sequence[Tensor | None]
| None = None,
audio_target_timesteps: Sequence[Tensor | None]
| None = None,
) -> dict[str, Any]
Build one DiT forward kwargs dict covering every request in the batch.
A single request delegates to :meth:MiniMaxH3DenoiseBranch.forward_kwargs, so step mode and request mode are the same call rather than two layouts that happen to agree -- which also keeps video_token_layout (the sparse- attention hint) on the single-request step path.