vllm_omni.diffusion.attention.backends.cudnn_attn ¶
CuDNNAttentionBackend ¶
Bases: AttentionBackend
CuDNNAttentionImpl ¶
Bases: AttentionImpl
backend_explicit instance-attribute ¶
backend_explicit = bool(
extra_impl_args.get("backend_explicit", False)
)
forward_cuda ¶
forward_cuda(
query: Tensor,
key: Tensor,
value: Tensor,
attn_metadata: AttentionMetadata | None = None,
) -> Tensor