Diffusion Attention Backend Selection¶
This document defines the selection and extension contract for diffusion attention backends. User-facing backend choices, installation, and tuning are in the attention backend guides.
Scope¶
The contract applies to DiT and other diffusion attention layers. It is separate from vLLM's autoregressive attention selector.
The selection path has four responsibilities:
- normalize user configuration into
AttentionConfigandAttentionSpec; - resolve a spec for an attention role;
- ask the active platform to validate an explicit backend or choose a hardware default; and
- load the selected
AttentionBackendclass from the registry path.
Resolution contract¶
get_attn_backend_for_role() resolves configuration in this order:
- exact
per_role[role]; - category
per_role[role_category]; default;- platform default.
An explicit resolution returns both the backend class and its AttentionSpec. A platform-default resolution returns the class and None. Layers must therefore treat the spec as optional and must not infer that a missing spec means SDPA.
Class resolution is cached by backend name, head size, and whether TRTLLM may be selected automatically. Logging is cached separately by role and source so the startup record identifies why each role received its backend.
Registry and platform boundary¶
DiffusionAttentionBackendEnum maps stable configuration names to default qualified class paths. register_diffusion_backend() may replace a path at runtime without changing the public enum value.
The active OmniPlatform owns compatibility policy through get_diffusion_attn_backend_cls(). It must:
- validate explicit selections and fail with an actionable error when the requested kernel cannot run;
- choose only an available, compatible default when no backend is explicit;
- account for head size and other platform-visible constraints; and
- return a qualified class path implementing
AttentionBackend.
The selector must not duplicate device capability or package-availability policy that belongs to the platform.
Typed backend options¶
Backend-specific settings remain typed fields on AttentionSpec rather than unstructured keyword dictionaries:
quantis consumed by FlashInfer and TRTLLM with backend-specific value validation;skip_softmaxis consumed by TRTLLM; andblock_sparseis consumed by block-sparse backends such as RainFusion; andfastvideo_vsa_topkis consumed by FastVideo VSA.
A backend reads only the fields it owns and rejects incompatible values. New options should be added to a shared typed spec only when more than one backend shares their semantics; otherwise add a dedicated typed block.
Model integration contract¶
Each diffusion attention call declares a stable role and, when appropriate, a broader category. Model-specific roles permit precise overrides; categories preserve common self or cross policies. A model also decides whether its path is compatible with automatic TRTLLM selection, because only the model knows whether masking and packing satisfy the kernel contract.
Model-owned implementation specializations¶
Use Attention(..., impl_overrides={backend_name: implementation_class}) when a model needs different computation within an existing backend contract. The model defines the implementation in its own directory and passes it when constructing the layer:
# MyModelVSAImpl is a model-owned subclass of FastVideoVSAImpl.
self.attention = Attention(
num_heads=num_heads,
head_size=head_size,
softmax_scale=head_size**-0.5,
causal=False,
role="self",
impl_overrides={"FASTVIDEO_VSA": MyModelVSAImpl},
)
The override is applied after role/platform selection, keyed by the selected backend's get_name(). It must subclass the selected implementation; incompatible platform implementations raise TypeError. Unselected overrides have no effect. Backend capabilities, constructor options, and shared parallel dispatch remain unchanged and must be honored by the specialization.
Use a new backend for a new provider/capability contract. custom_attention instead owns communication and requires skip_sequence_parallel=True.
Adding or changing a backend¶
An implementation change is complete only when it:
- implements the
AttentionBackendinterface and exposes a stable enum name; - adds platform validation and default routing where appropriate;
- defines and validates any typed configuration;
- tests explicit selection, default selection, and incompatible paths;
- documents installation, fallback behavior, and quality implications in the corresponding user guide; and
- updates the overview matrix without moving algorithm details back into the landing page.
FastVideo VSA model metadata¶
FastWan's text-to-video and image-to-video modes both use Wan22Pipeline. It operates on a flattened DiT sequence but partitions tokens in the original latent video grid. Wan integrations therefore attach the post-patch (T, H, W) shape as vsa_dit_seq_shape attention metadata. The backend validates that T * H * W equals the sequence length, derives the runtime block count, and selects the configured top-k key/value blocks for each query block.
Checkpoint-specific compensation remains model-owned. Wan layers pass an optional learned gate_compress projection through attention metadata when the checkpoint contains those weights. Checkpoints without the projection do not allocate or execute it. When top-k selects every block, native checkpoints route to SDPA, while FastVideo DMD checkpoints preserve the VSA all-block path to retain checkpoint semantics.
MiniMax-H3 supplies MiniMaxH3VSAImpl through impl_overrides. models/minimax_h3/attention/fastvideo_h3.py owns its prefix layout, video tiling, and learned compression gate. Pooling, block-map construction, and tile64 provider calls are shared through attention/ops/block_sparse.py.