vllm_omni.diffusion.distributed.autoencoders.wan_vae_fastpath.decode ¶
Chunked Wan decode with a preallocated output buffer.
decode_frames ¶
AutoencoderKLWan._decode's frame loop, writing every chunk straight into the result.
Upstream grows the result with torch.cat([out, out_], 2) on every chunk (quadratic copying) and then materializes unpatchify and clamp as two more full-size copies, so three or four copies of the decoded video are live at once. Here the final [B, C, T, H, W] buffer is allocated after the first chunk and each chunk is unpatchified and clamped directly into its slot. Every step is a permutation or an elementwise clamp, so the values are identical to upstream. Tiling dispatch is the caller's responsibility.