Skip to content

vllm_omni.diffusion.models.minimax_h3.attention.fastvideo_h3

MiniMax-H3 VSA prefix routing, tile layout, and learned compression gate.

logger module-attribute

logger = init_logger(__name__)

MiniMaxH3VSAImpl

Bases: FastVideoVSAImpl

Apply H3 tile64 routing after shared parallel attention dispatch.

forward_cuda

forward_cuda(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor