vllm_omni.diffusion.models.minimax_h3.npu ¶
MiniMax-H3 components for artifacts produced on Ascend NPU.
The package name records where these artifacts come from, not where they can run. Everything here is plain torch plus safetensors and carries no torch_npu dependency, so the same checkpoint loads and serves on CUDA and CPU as well. See tests/diffusion/models/minimax_h3/test_minimax_h3_native_lora.py for the platform-agnostic and CUDA coverage that pins this.
Modules:
| Name | Description |
|---|---|
lora | Loader for the native-layout MiniMax-H3 distilled LoRA (FlashGen v1.0). |
load_minimax_h3_native_lora ¶
load_minimax_h3_native_lora(
*,
partition: str,
lora_request: LoRARequest,
lora_path: str | Path,
dtype: dtype,
unsupported_offload_mode: str | None = None,
) -> tuple[LoRAModel, PEFTHelper, DMD2SigmaSchedule] | None
Load a native-layout MiniMax-H3 distilled LoRA through the legacy manager.