Skip to content

vllm_omni.diffusion.layers.fused_qk_rope

Bit-exact paired Q/K full-width interleaved RoPE for CUDA.

The operator accepts already-normalized BF16 query and key tensors in BSND layout and full-width FP32 cosine/sine tables. It deliberately performs two round-to-nearest FP32 multiplies followed by one round-to-nearest FP32 add, then rounds only once when storing BF16 output. This matches Diffusers' apply_rotary_emb arithmetic while combining its separate Q and K calls.

fused_qk_rope

fused_qk_rope(
    q: Tensor, k: Tensor, cos: Tensor, sin: Tensor
) -> tuple[Tensor, Tensor]

Apply paired full-width adjacent/interleaved RoPE on CUDA.

fused_qk_rope_supported

fused_qk_rope_supported(
    q: Tensor, k: Tensor, cos: Tensor, sin: Tensor
) -> bool

Return whether the exact CUDA kernel can consume these tensors.