vllm_omni.diffusion.attention.backends.flash_attn ¶
FlashAttentionBackend ¶
Bases: AttentionBackend
supports_attention_mask classmethod ¶
FlashAttentionImpl ¶
Bases: AttentionImpl[AttentionMetadata]
fa_deterministic instance-attribute ¶
forward_cuda ¶
forward_cuda(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
CUDA/ROCm/MUSA flash attention implementation.
forward_fa_npu ¶
forward_fa_npu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
forward_fa_quant_npu ¶
forward_fa_quant_npu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
forward_npu ¶
forward_npu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
NPU attention implementation using mindiesd.
forward_paged ¶
Run platform-native paged attention through the Omni backend.
DiffusionPagedAttentionAdapter owns BlockTables and prepares the rank-local native cache context. The adapter no longer performs the attention call itself; this method is the backend-owned execution boundary. Its native layer wrapper keeps vLLM version-specific cache and kernel details out of Omni's common Attention layer. CUDA uses the native vLLM FlashAttention writer/kernel contract. Ascend writes the complete K/V span once before its piecewise FIA calls.
forward_xpu ¶
forward_xpu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
XPU flash attention implementation.