vllm_omni.diffusion.models.minimax_h3.npu.lora ¶
Loader for the native-layout MiniMax-H3 distilled LoRA (FlashGen v1.0).
This contract is keyed by key_format=minimax-h3-native and is distinct from the LightX2V Turbo contract in the sibling ..lora module: it keeps the fused qkv_proj, adds adaln_proj.linear, uses rank 64, and declares its own rectified-flow schedule in safetensors metadata.
The artifact is trained on Ascend NPU, but nothing below is NPU-specific. Only torch and safetensors are used, tensors are read on CPU, and the compute device is chosen later by the pipeline, so the same file serves on CUDA too.
load_minimax_h3_native_lora ¶
load_minimax_h3_native_lora(
*,
partition: str,
lora_request: LoRARequest,
lora_path: str | Path,
dtype: dtype,
unsupported_offload_mode: str | None = None,
) -> tuple[LoRAModel, PEFTHelper, DMD2SigmaSchedule] | None
Load a native-layout MiniMax-H3 distilled LoRA through the legacy manager.