Skip to content

vllm_omni.diffusion.models.minimax_h3.continuation

Bounded H3 windows with synchronized latent-tail guides and global AV RoPE positions.

The guide/discard/append algorithm follows ComfyUI-Minimax-H3-Continuation: https://github.com/ttulttul/ComfyUI-Minimax-H3-Continuation

logger module-attribute

logger = init_logger(__name__)

ContinuationWindow dataclass

audio_end property

audio_end: int

audio_start property

audio_start: int

end instance-attribute

end: int

overlap instance-attribute

overlap: int

overlap_audio_t property

overlap_audio_t: int

start instance-attribute

start: int

diffuse_continuation

diffuse_continuation(
    diffuse: Callable[..., tuple[Tensor, Tensor]],
    kwargs: dict[str, Any],
    *,
    window_frames: int,
    overlap_frames: int,
    text_conditioning: Sequence[tuple[Tensor, Tensor]]
    | None = None,
) -> tuple[Tensor, Tensor]

Denoise fresh windows; retain old latents and append only new suffixes.

Guides are extra condition rows sharing the new target's temporal origin, not a masked target prefix. Audio boundaries refer to the cumulative frame timeline, preventing per-window rounding from accumulating A/V drift. Each window shifts temporal media positions onto that same global timeline before RoPE is evaluated; text and static image references remain fixed.

plan_continuation_windows

plan_continuation_windows(
    total_frames: int,
    window_frames: int,
    overlap_frames: int,
) -> list[ContinuationWindow]

resolve_continuation

resolve_continuation(
    extra: Mapping[str, Any],
    *,
    task: str,
    step_execution: bool,
) -> tuple[int, int] | None