vllm_omni.diffusion.layers.fused_qk_rope ¶
Bit-exact paired Q/K full-width interleaved RoPE for CUDA.
The operator accepts already-normalized BF16 query and key tensors in BSND layout and full-width FP32 cosine/sine tables. It deliberately performs two round-to-nearest FP32 multiplies followed by one round-to-nearest FP32 add, then rounds only once when storing BF16 output. This matches Diffusers' apply_rotary_emb arithmetic while combining its separate Q and K calls.