vllm_omni.platforms.npu.platform ¶
NPUOmniPlatform ¶
Bases: OmniPlatform, NPUPlatform
NPU/Ascend implementation of OmniPlatform.
Inherits all NPU-specific implementations from vllm-ascend's NPUPlatform, and adds Omni-specific interfaces from OmniPlatform.
build_diffusion_kv_attn_metadata classmethod ¶
Build the Ascend metadata required by the native NPU backend.
configure_diffusion_vllm_config classmethod ¶
Use the block geometry required by Ascend's native paged kernel.
create_autocast_context classmethod ¶
get_device_memory classmethod ¶
get_diffusion_attn_backend_cls classmethod ¶
get_diffusion_attn_backend_cls(
selected_backend: str | None,
head_size: int,
allow_trtllm_default: bool = True,
) -> str
get_diffusion_packed_modules_mapping classmethod ¶
get_diffusion_paged_kv_attn_backend classmethod ¶
Keep strict Ulysses paged FIA out of vLLM's PCP implementation.
init_diffusion_model_runner_runtime classmethod ¶
init_diffusion_worker_vllm_config classmethod ¶
init_diffusion_worker_vllm_config(vllm_config: Any) -> None
prepare_diffusion_op_runtime classmethod ¶
record_device_event classmethod ¶
Record a NPU event on the default stream to mark tensor readiness.
On NPU/Ascend with HCCL, distributed communication may use internal streams not visible to the default stream. Synchronize the default stream first so that HCCL results are written back before we record the event, ensuring d2h_stream.wait_event() captures the complete output data.
register_additional_diffusion_fused_moe_hooks classmethod ¶
register_additional_diffusion_fused_moe_hooks(
moe_runner: Any,
) -> None
requires_diffusion_paged_kv_prewrite classmethod ¶
requires_diffusion_paged_kv_prewrite() -> bool
Write the full K/V span once before piecewise FIA segments.