vllm_omni.platforms.cuda ¶
Modules:
| Name | Description |
|---|---|
platform | |
CudaOmniPlatform ¶
Bases: OmniPlatform, CudaPlatformBase
CUDA/GPU implementation of OmniPlatform (default).
Inherits all CUDA-specific implementations from vLLM's CudaPlatform, and adds Omni-specific interfaces from OmniPlatform.
get_default_ir_op_priority classmethod ¶
Prefer vllm_c CUDA kernels over native for diffusion IR ops.
get_device_capability classmethod ¶
get_device_capability(
device_id: int = 0,
) -> DeviceCapability | None
get_device_memory classmethod ¶
get_diffusion_attn_backend_cls classmethod ¶
get_diffusion_attn_backend_cls(
selected_backend: str | None,
head_size: int,
allow_trtllm_default: bool = True,
) -> str
has_flash_attn_4 classmethod ¶
has_flash_attn_4() -> bool
Return whether CuTe FA4 is importable (Blackwell-capable FLASH_ATTN).
record_device_event classmethod ¶
Record a device event on the default stream to mark tensor readiness.