vllm_omni.diffusion.attention.backends.rainfusion_attn ¶
RainFusionAttentionBackend ¶
Bases: AttentionBackend
RainFusionAttentionImpl ¶
Bases: AttentionImpl
Block-sparse video attention via MindIE-SD RainFusion (rf_v2) on Ascend NPU.
Sparsity applies only to the video segment of a packed multimodal sequence, whose extent the model publishes as AttentionMetadata.video_layout. Every other case — warmup denoise steps, exempt layers, a layer that does not declare qkv_layout="BSND", sequences without a published video segment, video segments too short to pay for block selection — delegates to FlashAttention, so a model can select this backend unconditionally. MindIE-SD handles an irregular video tail internally, retaining it outside the sparse blocks.
dense_fallback instance-attribute ¶
dense_fallback = FlashAttentionBackend.get_impl_cls()(
num_heads=num_heads,
head_size=head_size,
softmax_scale=softmax_scale,
causal=causal,
num_kv_heads=num_kv_heads,
prefix=prefix,
qkv_layout=qkv_layout,
)
forward_cuda ¶
forward_cuda(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
forward_npu ¶
forward_npu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
forward_xpu ¶
forward_xpu(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor
RainFusionConfig dataclass ¶
Resolved RainFusion controls for one attention layer.
sparsity is the nominal fraction of key blocks dropped per query block. The realized sparsity is lower because rf_v2 always keeps the prefix rows and the first-frame blocks. start_step and skip_layers are the accuracy knobs: early denoise steps and specific DiT blocks stay dense.
from_backend_kwargs classmethod ¶
from_backend_kwargs(
backend_kwargs: dict | None,
) -> RainFusionConfig
RainFusionPlan dataclass ¶
Per-forward geometry handed to the rf_v2 kernel.