Skip to content

vllm_omni.diffusion.attention.backends.flash_attn

logger module-attribute

logger = init_logger(__name__)

FlashAttentionBackend

Bases: AttentionBackend

accept_output_buffer class-attribute instance-attribute

accept_output_buffer: bool = True

supports_paged_kv class-attribute instance-attribute

supports_paged_kv: bool = True

supports_piecewise_spans class-attribute instance-attribute

supports_piecewise_spans: bool = True

get_impl_cls staticmethod

get_impl_cls() -> type[FlashAttentionImpl]

get_name staticmethod

get_name() -> str

get_supported_head_sizes staticmethod

get_supported_head_sizes() -> list[int]

supports_attention_mask classmethod

supports_attention_mask(
    attention_spec: object | None = None,
) -> bool

supports_multi_doc_packed_varlen classmethod

supports_multi_doc_packed_varlen() -> bool

supports_packed_mask_free classmethod

supports_packed_mask_free() -> bool

FlashAttentionImpl

Bases: AttentionImpl[AttentionMetadata]

causal instance-attribute

causal = causal

fa_deterministic instance-attribute

fa_deterministic = (
    bool(getattr(cfg, "fa_deterministic", False))
    if cfg is not None
    else False
)

is_cross_attn instance-attribute

is_cross_attn = role == 'cross'

num_heads instance-attribute

num_heads = num_heads

qkv_layout instance-attribute

qkv_layout = qkv_layout

softmax_scale instance-attribute

softmax_scale = softmax_scale

forward_cuda

forward_cuda(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor

CUDA/ROCm/MUSA flash attention implementation.

forward_fa_npu

forward_fa_npu(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor

forward_fa_quant_npu

forward_fa_quant_npu(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor

forward_npu

forward_npu(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor

NPU attention implementation using mindiesd.

forward_paged

forward_paged(paged_kv_context) -> Tensor

Run platform-native paged attention through the Omni backend.

DiffusionPagedAttentionAdapter owns BlockTables and prepares the rank-local native cache context. The adapter no longer performs the attention call itself; this method is the backend-owned execution boundary. Its native layer wrapper keeps vLLM version-specific cache and kernel details out of Omni's common Attention layer. CUDA uses the native vLLM FlashAttention writer/kernel contract. Ascend writes the complete K/V span once before its piecewise FIA calls.

forward_xpu

forward_xpu(
    query: Tensor,
    key: Tensor,
    value: Tensor,
    attn_metadata: AttentionMetadata | None = None,
) -> Tensor

XPU flash attention implementation.