Skip to content

vllm_omni.diffusion.models.minimax_h3.condition_noise

MiniMax H3 visual/audio condition-noise augmentation.

The request's condition timestep is applied to both the tensor value and the DiT timestep. Tokenizer artifacts remain clean and reusable; this module materializes the fixed noised anchors immediately before the denoise loop.

MINIMAX_H3_AUDIO_COND_CHANNELS module-attribute

MINIMAX_H3_AUDIO_COND_CHANNELS = 2

minimax_h3_audio_cond_noise_aug_rows

minimax_h3_audio_cond_noise_aug_rows(
    clean_rows: Tensor,
    *,
    condition_audio_t: Sequence[int],
    seed: int,
    noise_aug: float,
) -> Tensor

Apply the audio-condition RF noise recipe to packed clean rows.

condition_audio_t contains the latent T of each audio-bearing condition in canonical request order. Noise is drawn per condition element, with a fresh CPU generator seeded with seed + 1 for every element. Consequently each condition restarts the same RNG stream; concatenating the rows and drawing once would be numerically different for ordered multi-reference requests.

The mix is intentionally evaluated on CPU in fp32 before the packed rows are transferred to the DiT device.

minimax_h3_imgvid_cond_noise_aug_rows

minimax_h3_imgvid_cond_noise_aug_rows(
    clean_rows: Tensor,
    *,
    condition_shapes: Sequence[Sequence[int]],
    target_latent_t: int,
    imgvid_cond_num_frames: int,
    seed: int,
    noise_aug: float,
) -> Tensor

Apply the imgvid-condition RF noise recipe to packed clean rows.

condition_shapes contains (latent_t, latent_h, latent_w) in packed visual-condition order. A new CPU generator with the same row seed is created for every condition. Under the dependent-noise policy, each draw uses the target temporal length plus the template's imgvid-condition frame count, then slices the prefix matching the current condition.