Skip to content

vllm_omni.diffusion.models.minimax_h3.latent_mask

MiniMax H3 latent-edit mask parsing and sampler math.

MiniMaxH3LatentEdit dataclass

Clean source, model anchor, and masks for one packed latent stream.

anchor_rows instance-attribute

anchor_rows: Tensor

clean_rows instance-attribute

clean_rows: Tensor

mask_rows instance-attribute

mask_rows: Tensor

restore_mask_rows instance-attribute

restore_mask_rows: Tensor

from_rows classmethod

from_rows(
    clean_rows: Tensor,
    anchor_rows: Tensor,
    mask_rows: Tensor,
    restore_mask_rows: Tensor,
) -> MiniMaxH3LatentEdit

Create an edit from already-validated encoder outputs.

Value validation and all-generate elision belong at the encoder input boundary. This constructor only normalizes devices/dtypes and checks the cheap row-shape invariants needed by the sampler math.

model_rows

model_rows(state_rows: Tensor) -> Tensor

prepare

prepare(
    state_rows: Tensor,
    timestep: float,
    condition_timestep: float,
    *,
    sigma: float,
) -> tuple[Tensor, Tensor]

Prepare the model rows and per-row timesteps for one denoise step.

target_timesteps

target_timesteps(
    timestep: float,
    condition_timestep: float,
    *,
    sigma: float | None = None,
) -> Tensor

to

to(
    *, device: device, dtype: dtype = float32
) -> MiniMaxH3LatentEdit

x0

x0(
    model_rows: Tensor, velocity: Tensor, timestep: float
) -> Tensor

MiniMaxH3ParsedMask dataclass

Token mask for the model and raw mask for the final x0 restore.

model_mask_rows instance-attribute

model_mask_rows: Tensor

restore_mask_rows instance-attribute

restore_mask_rows: Tensor

minimax_h3_audio_edit_masks

minimax_h3_audio_edit_masks(
    mask: Tensor, *, audio_t: int
) -> MiniMaxH3ParsedMask

Flatten the canonical channel-major audio mask to H3 row order.

minimax_h3_prepare_edit_rows

minimax_h3_prepare_edit_rows(
    rows: Tensor,
    update_mask: Tensor,
    edit: MiniMaxH3LatentEdit | None,
    timestep: float,
    condition_timestep: float,
    *,
    sigma: float,
) -> tuple[Tensor, Tensor | None]

Build one model input view without mutating the persistent sampler rows.

minimax_h3_video_edit_masks

minimax_h3_video_edit_masks(
    mask: Tensor,
    *,
    latent_t: int,
    latent_h: int,
    latent_w: int,
) -> MiniMaxH3ParsedMask

Map the canonical full-grid video mask to model and restore rows.

Full-grid masks are max-pooled over each 2x2 DiT token. Their raw four cells are repeated in the packed 24-channel feature order for restoration.