Skip to content

vllm_omni.platforms.npu.quant.kv_quant_npu

Seeded rotation matrices for MindIE-SD quantized attention.

get_quant_attention_rotation cached

get_quant_attention_rotation(
    device: device, dtype: dtype, head_dim: int
) -> Tensor

Preserve Omni's fixed Walsh-Hadamard rotation without changing global RNG.