Skip to content

vllm_omni.diffusion.models.minimax_h3.npu

MiniMax-H3 components for artifacts produced on Ascend NPU.

The package name records where these artifacts come from, not where they can run. Everything here is plain torch plus safetensors and carries no torch_npu dependency, so the same checkpoint loads and serves on CUDA and CPU as well. See tests/diffusion/models/minimax_h3/test_minimax_h3_native_lora.py for the platform-agnostic and CUDA coverage that pins this.

Modules:

Name Description
lora

Loader for the native-layout MiniMax-H3 distilled LoRA (FlashGen v1.0).

MINIMAX_H3_NATIVE_INFERENCE_STEPS module-attribute

MINIMAX_H3_NATIVE_INFERENCE_STEPS = 4

load_minimax_h3_native_lora

load_minimax_h3_native_lora(
    *,
    partition: str,
    lora_request: LoRARequest,
    lora_path: str | Path,
    dtype: dtype,
    unsupported_offload_mode: str | None = None,
) -> tuple[LoRAModel, PEFTHelper, DMD2SigmaSchedule] | None

Load a native-layout MiniMax-H3 distilled LoRA through the legacy manager.