vllm_omni.diffusion.models.minimax_h3.vae_temporal ¶
Memory-optimized forks of the checkpoint's temporal VAE loops.
Two method replacements installed on the checkpoint's AutoencoderKLLegacy at adapter construction:
-
encode_temporalpads misaligned frame counts by concatenating repeated last frames onto the whole video before the chunk loop (~4GB copy for a 15s clip). The fork pads inside the tail chunk instead: the chunks see bit-identical frames while the copy shrinks to a single clip. -
_decode_temporal_streamingaccumulates the whole decoded video in a floating-point buffer that is denormalized, clamped, and quantized to uint8 only after the final chunk lands. The fork fuses the adapter's in-place revert and the output quantizer intowrite_partand accumulates directly into a uint8 buffer -- the same per-element op order ((x - mean) / std -> clamp(0, 1) -> *255 -> round -> uint8), so the delivered video is bit-identical while the resident decode buffer shrinks 4x.
Set VLLM_OMNI_VAE_LEGACY_TEMPORAL=1 to keep the checkpoint's methods.
install_temporal_stream_patches ¶
Install the memory-optimized temporal forks on the checkpoint model.
Each fork is installed only when the checkpoint contract it mirrors is discoverable; otherwise the checkpoint's own method keeps running (the adapter's post-decode revert already covers that path).