Skip to content

vllm_omni.diffusion.models.minimax_h3.fasth3

FastVideo FastH3: a four-step DMD2 student of MiniMax-H3.

FastH3 replaces H3's 49 denoiser evaluations with four. It ships as an adapter over the base checkpoint rather than as a full release, so it reuses H3's text encoder, video VAE, audio VAE, tokenizers and schedulers unchanged.

The artifact is not a PEFT LoRA, and it is not request-switchable. Its own metadata states the reconstruction as::

W = W_base + lora_B @ lora_A; then .diff/.diff_b added and .set_weight assigned

so besides rank-64 factors it carries full-rank .diff/.diff_b deltas for RMSNorm weights, biases, patch projections and the final layer - none of which a LoRA layer can express - and the VSA variants add .set_weight tensors for compression gates that do not exist in the base transformer at all. The adapter is therefore fused into the checkpoint stream at load time, before the weights are sharded, which is also what the release's model card requires.

The low-rank factors carry no alpha: the reconstruction adds lora_B @ lora_A directly, i.e. a scale of exactly 1.

Two checkpoint spellings meet here. The adapter is written in the diffusers namespace (transformer_blocks.0.attn.to_q) while vLLM-Omni loads H3's native one (blocks.0.attn.qkv_proj), whose attention and MLP projections are fused. Every mapping and layout convention below was verified tensor by tensor against the released full checkpoint (FastVideo-FastH3-4-step-Preview-v1-Dense-DataFree): W_base + delta reproduces it to bf16 rounding.

FASTH3_BASE_MODEL module-attribute

FASTH3_BASE_MODEL = 'MiniMaxAI/MiniMax-H3'

FASTH3_BASE_SCHEDULE module-attribute

FASTH3_BASE_SCHEDULE = DMD2SigmaSchedule.from_positions(
    (0.999, 0.749, 0.5, 0.25, 0.0)
)

FASTH3_DENOISE_STEPS module-attribute

FASTH3_DENOISE_STEPS = (
    FASTH3_BASE_SCHEDULE.num_inference_steps
)

FASTH3_FORMAT module-attribute

FASTH3_FORMAT = 'fastvideo-lora-v2'

FASTH3_MANIFEST module-attribute

FASTH3_MANIFEST = 'adapter_manifest.json'

FASTH3_SUPPORTED_TASKS module-attribute

FASTH3_SUPPORTED_TASKS = frozenset({'t2va'})

logger module-attribute

logger = init_logger(__name__)

FastH3AdapterError

Bases: ValueError

The artifact is a FastH3 adapter, but it cannot be applied as one.

FastH3WeightFusion

Fuse a FastH3 adapter into the H3 checkpoint stream as it is loaded.

base_schedule property

base_schedule: tuple[float, ...]

The rectified-flow positions this student samples on.

The fused checkpoint is a four-step student, so the ladder comes from the release rather than from the many-step teacher's metadata or from the uniform one num_inference_steps would otherwise derive.

requires_vsa instance-attribute

requires_vsa = requires_vsa

source property

source: Path

apply

apply(
    weights: Iterable[tuple[str, Tensor]],
) -> Iterator[tuple[str, Tensor]]

Fuse every streamed checkpoint tensor on its way into the model.

check_request

check_request(
    sampling: Any, *, video_shift: float, audio_shift: float
) -> None

Refuse a request that would sample the student off its rungs.

check_serving_contract

check_serving_contract(
    *,
    partition: str,
    od_config: Any,
    video_shift: float,
    audio_shift: float,
) -> None

Hold a starting server to the ladder this student was trained on.

check_task

check_task(task: str) -> None

Refuse a task this preview never distilled.

from_path classmethod

from_path(
    path: str | Path,
    *,
    head_dim: int,
    num_blocks: int,
    num_refiner_blocks: int,
) -> FastH3WeightFusion | None

Build a fusion from an adapter file or directory, else None.

Returning None keeps every other --lora-path artifact on the dynamic LoRA route; only a file carrying the FastH3 release identity is claimed here.

The block counts are the model's, and the artifact has to cover them: claiming a partial adapter would switch the server onto the four-step contract while most of the transformer still held base H3 weights.

fuse

fuse(name: str, weight: Tensor) -> Tensor

Return weight with this adapter's contribution added.

validate_fully_applied

validate_fully_applied(
    loaded: Iterable[str] | None = None,
) -> None

Close the fusion: every edit must have met its parameter.

A silently unapplied delta is the failure mode that matters here: the model would load and generate, just not as the distilled student. The weights are loaded once, so the mapped payloads are dropped afterwards rather than held for the life of the process.

loaded is the set of parameter names load_weights actually consumed. A gate is assigned rather than fused, so it lands on a module the base transformer does not have; if that module was never built, load_weights only logs a skip and the server would serve a zero-initialized gate. Yielding a tensor is not evidence it arrived, so the injections are closed against that set when it is available.

resolve_fasth3_fusion

resolve_fasth3_fusion(
    od_config: Any, transformer: MiniMaxH3DiTModel
) -> FastH3WeightFusion | None

Claim --lora-path when it points at a FastH3 adapter.

FastH3 rewrites RMSNorm weights and biases, so it cannot be expressed as a request-switchable LoRA and is fused into the checkpoint instead. Any other artifact returns None here and stays on the dynamic LoRA route.