Skip to content

vllm_omni.platforms.npu.models.step_audio2_token2wav

NPU patches for Step-Audio2 / MiniCPM Token2Wav.

Ascend-specific workarounds that must not live in the shared GPU model file:

  1. HiFT sine-source downsample — replace the failing 480x linear1d downsample with its exact midpoint form while keeping HiFT on NPU.
  2. CosyVoice2 DiT SDPA — force MATH backend (+ DiT attn mask expand) to avoid fused FA rejecting CosyVoice (B,1,1,S) masks (error 161001).

logger module-attribute

logger = init_logger(__name__)

apply_step_audio2_token2wav_npu_patch

apply_step_audio2_token2wav_npu_patch() -> None

Monkey-patch StepAudio2Token2WavCore for Ascend NPU.

Import is deferred and optional: platform bootstrap (e.g. resolving current_omni_platform from rotary embedding) must not require Token2Wav optional deps such as librosa.

npu_token2wav_sdpa_context

npu_token2wav_sdpa_context(
    *, require_math: bool = False
) -> Iterator[None]

Expand CosyVoice masks + force MATH SDPA to avoid FA 161001.

patch_step_audio2_hift_for_npu

patch_step_audio2_hift_for_npu(hift: Module) -> None

Patch the non-causal Step-Audio2 HiFT implementation for Ascend.

The flashcosyvoice.SineGen2 instantiated by Step-Audio2 1.0.0 is non-causal and reduces a full-rate phase tensor by 1 / 480 before restoring it to the waveform rate. Ascend's upsample_linear1d kernel can raise an AIVector UB-address exception (ACL 507015) for that reduction.

The exact midpoint form keeps the common path on NPU. Unsupported or pulse configurations delegate only _f02sine to CPU, preserving upstream behavior without restoring the old whole-HiFT CPU offload.