Skip to content

vllm_omni.diffusion.layers.fused_qk_norm_rope

Q/K RMSNorm followed by packed RoPE (non-interleaved or interleaved).

The public contract is shared by diffusion attention implementations:

  • q and k are [tokens, heads, head_dim];
  • norm weights are one-dimensional [head_dim] tensors;
  • rope_table is [tokens, rotary_dim] and stores [cos(theta), sin(theta)] with theta of width rotary_dim // 2; its dtype is either the activation dtype or float32;
  • interleaved=False rotates half-split pairs (d, d + rotary_dim/2) (MiniMax-H3); interleaved=True rotates adjacent pairs (2i, 2i + 1) with theta_i (Boogu-Image, matching its apply_rotary_emb).

The CUDA fast path fuses RMSNorm and RoPE without materializing normalized Q/K or rotary-product intermediates. The interleaved mode supports any even rotary_dim <= head_dim <= 256 (tl.arange padding to the next power of two); the half-split mode keeps the pre-existing kernel and its MiniMax-H3 geometry contract (head_dim == 128, rotary_dim == 96). Ascend composes its RMSNorm and rotary fused primitives on the MiniMax-H3 geometry; unsupported inputs use the eager reference.

fused_qk_norm_rope

fused_qk_norm_rope(
    q: Tensor,
    k: Tensor,
    q_weight: Tensor,
    k_weight: Tensor,
    rope_table: Tensor,
    eps: float,
    *,
    head_dim: int | None = None,
    rotary_dim: int | None = None,
    interleaved: bool = False,
) -> tuple[Tensor, Tensor]

Apply Q/K RMSNorm and packed RoPE (half-split or adjacent pairs).

fused_qk_norm_rope_min_tokens

fused_qk_norm_rope_min_tokens(default: int) -> int

Token span below which a consumer should keep its eager chain.

Taking the fused path costs a fixed amount of host time per call (the custom-op dispatch into the Python impl and the Triton launcher), while the GPU time it saves grows with the number of tokens the call covers. Below a host- and stack-dependent crossover a call is host-bound and the fused path can cost more end-to-end than it saves; above it the fused path wins. Consumers pass the default they measured on their reference hardware (Boogu-Image: 2048 tokens on one H200) and a deployment on other hardware overrides it once, for every consumer, through VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS (0 = always fuse).

Resolved at call time — consumers call this once per forward, not per attention site — so tests and operators can flip it without re-importing. Unset or blank means the caller's default; anything else that is not a non-negative integer raises ValueError.