vllm_omni.diffusion.models.minimax_h3.condition_noise ¶
MiniMax H3 visual/audio condition-noise augmentation.
The request's condition timestep is applied to both the tensor value and the DiT timestep. Tokenizer artifacts remain clean and reusable; this module materializes the fixed noised anchors immediately before the denoise loop.
minimax_h3_audio_cond_noise_aug_rows ¶
minimax_h3_audio_cond_noise_aug_rows(
clean_rows: Tensor,
*,
condition_audio_t: Sequence[int],
seed: int,
noise_aug: float,
) -> Tensor
Apply the audio-condition RF noise recipe to packed clean rows.
condition_audio_t contains the latent T of each audio-bearing condition in canonical request order. Noise is drawn per condition element, with a fresh CPU generator seeded with seed + 1 for every element. Consequently each condition restarts the same RNG stream; concatenating the rows and drawing once would be numerically different for ordered multi-reference requests.
The mix is intentionally evaluated on CPU in fp32 before the packed rows are transferred to the DiT device.
minimax_h3_imgvid_cond_noise_aug_rows ¶
minimax_h3_imgvid_cond_noise_aug_rows(
clean_rows: Tensor,
*,
condition_shapes: Sequence[Sequence[int]],
target_latent_t: int,
imgvid_cond_num_frames: int,
seed: int,
noise_aug: float,
) -> Tensor
Apply the imgvid-condition RF noise recipe to packed clean rows.
condition_shapes contains (latent_t, latent_h, latent_w) in packed visual-condition order. A new CPU generator with the same row seed is created for every condition. Under the dependent-noise policy, each draw uses the target temporal length plus the template's imgvid-condition frame count, then slices the prefix matching the current condition.