Skip to content

vllm_omni.diffusion.models.minimax_h3.vae_temporal

Memory-optimized forks of the checkpoint's temporal VAE loops.

Two method replacements installed on the checkpoint's AutoencoderKLLegacy at adapter construction:

  • encode_temporal pads misaligned frame counts by concatenating repeated last frames onto the whole video before the chunk loop (~4GB copy for a 15s clip). The fork pads inside the tail chunk instead: the chunks see bit-identical frames while the copy shrinks to a single clip.

  • _decode_temporal_streaming accumulates the whole decoded video in a floating-point buffer that is denormalized, clamped, and quantized to uint8 only after the final chunk lands. The fork fuses the adapter's in-place revert and the output quantizer into write_part and accumulates directly into a uint8 buffer -- the same per-element op order ((x - mean) / std -> clamp(0, 1) -> *255 -> round -> uint8), so the delivered video is bit-identical while the resident decode buffer shrinks 4x.

Set VLLM_OMNI_VAE_LEGACY_TEMPORAL=1 to keep the checkpoint's methods.

logger module-attribute

logger = init_logger(__name__)

install_temporal_stream_patches

install_temporal_stream_patches(model) -> None

Install the memory-optimized temporal forks on the checkpoint model.

Each fork is installed only when the checkpoint contract it mirrors is discoverable; otherwise the checkpoint's own method keeps running (the adapter's post-decode revert already covers that path).