Skip to content

vllm_omni.platforms.npu.models.minimax_h3

NPU patches for the MiniMax H3 Qwen3-VL text encoder.

logger module-attribute

logger = init_logger(__name__)

apply_minimax_h3_qwen3vl_patch

apply_minimax_h3_qwen3vl_patch() -> None

Route MiniMax H3 Qwen3-VL text RoPE to the Ascend fused operator.

apply_minimax_h3_qwen3vl_sdpa_patch

apply_minimax_h3_qwen3vl_sdpa_patch() -> None

Route MiniMax H3 Qwen3-VL text attention to NPU native GQA.

apply_minimax_h3_qwen3vl_swiglu_patch

apply_minimax_h3_qwen3vl_swiglu_patch() -> None

Route MiniMax H3 Qwen3-VL text MLP to the Ascend fused SwiGLU path.

npu_swiglu_from_packed

npu_swiglu_from_packed(gate_up: Tensor) -> Tensor

Apply fused SwiGLU to a packed gate/up projection tensor.