Skip to content

vllm_omni.diffusion.models.minimax_h3.npu.lora

Loader for the native-layout MiniMax-H3 distilled LoRA (FlashGen v1.0).

This contract is keyed by key_format=minimax-h3-native and is distinct from the LightX2V Turbo contract in the sibling ..lora module: it keeps the fused qkv_proj, adds adaln_proj.linear, uses rank 64, and declares its own rectified-flow schedule in safetensors metadata.

The artifact is trained on Ascend NPU, but nothing below is NPU-specific. Only torch and safetensors are used, tensors are read on CPU, and the compute device is chosen later by the pipeline, so the same file serves on CUDA too.

MINIMAX_H3_NATIVE_INFERENCE_STEPS module-attribute

MINIMAX_H3_NATIVE_INFERENCE_STEPS = 4

load_minimax_h3_native_lora

load_minimax_h3_native_lora(
    *,
    partition: str,
    lora_request: LoRARequest,
    lora_path: str | Path,
    dtype: dtype,
    unsupported_offload_mode: str | None = None,
) -> tuple[LoRAModel, PEFTHelper, DMD2SigmaSchedule] | None

Load a native-layout MiniMax-H3 distilled LoRA through the legacy manager.