Skip to content

vllm_omni.diffusion.models.minimax_h3.vae

MiniMax H3 remote-code VAE adapters and exact latent contracts.

MINIMAX_H3_AUDIO_CHANNELS module-attribute

MINIMAX_H3_AUDIO_CHANNELS = 2

MINIMAX_H3_AUDIO_SAMPLE_RATE module-attribute

MINIMAX_H3_AUDIO_SAMPLE_RATE = 32000

MINIMAX_H3_KEYFRAME_ENCODE_SEED module-attribute

MINIMAX_H3_KEYFRAME_ENCODE_SEED = 42

logger module-attribute

logger = init_logger(__name__)

MiniMaxH3AudioVAE

Bases: Module

config_dict instance-attribute

config_dict = _load_component_config(component_path)

decode_only instance-attribute

decode_only = bool(decode_only)

encode_only instance-attribute

encode_only = bool(encode_only)

model instance-attribute

model = self.remote.model

remote instance-attribute

remote = _load_audio_vae_encoder(
    component_path, self.config_dict
)

sample_rate instance-attribute

sample_rate = int(self.config_dict['sample_rate'])

decode_latent

decode_latent(latent: Tensor) -> Tensor

encode_waveform

encode_waveform(
    waveform: Tensor, sample_rate: int
) -> tuple[Tensor, int]

load_to_device

load_to_device() -> None

offload_to_cpu

offload_to_cpu() -> None

set_omni_component_cache

set_omni_component_cache(
    cache: BoundedAllocatorCache | None,
) -> None

MiniMaxH3VideoVAE

Bases: Module, DistributedVaeMixin

Adapter around the checkpoint's native parallel-tiled video VAE.

chunk_value_range class-attribute

chunk_value_range: tuple[float, float] = (0.0, 1.0)

config_dict instance-attribute

config_dict = _load_component_config(component_path)

decode_only instance-attribute

decode_only = bool(decode_only)

decoder_component instance-attribute

decoder_component = _VideoVAEPartProxy(self, 'decoder')

device_module instance-attribute

device_module = torch.get_device_module()

encode_only instance-attribute

encode_only = bool(encode_only)

encoder_component instance-attribute

encoder_component = _VideoVAEPartProxy(self, 'encoder')

model instance-attribute

model = self.remote.model

parallel_size instance-attribute

parallel_size = 1

remote instance-attribute

remote = _load_video_vae_encoder(
    component_path, self.config_dict
)

use_slicing instance-attribute

use_slicing = False

use_tiling instance-attribute

use_tiling = True

decode_latent

decode_latent(latent: Tensor) -> Tensor

decode_with_chunks

decode_with_chunks(
    z: Tensor, *, on_chunk: DecodedChunkConsumer
) -> None

Decode temporal clips and synchronously publish frames-only chunks.

Implements :class:SupportsChunkedVAEDecode. Every rank participating in distributed VAE execution must invoke this method with a callback so the temporal collectives stay in lockstep; on_chunk is called only on the rank that owns output. Chunks arrive as [B, C, T, H, W] float frames, normalized through the checkpoint's processor to match the complete decode path. After a callback failure, the remaining chunks are decoded and discarded before the exception is re-raised.

encode_image

encode_image(image: Image) -> Tensor

encode_video

encode_video(
    frames: Any,
) -> tuple[Tensor, tuple[int, int, int]]

is_distributed_enabled

is_distributed_enabled() -> bool

load_to_device

load_to_device() -> None

offload_to_cpu

offload_to_cpu() -> None

set_omni_component_cache

set_omni_component_cache(
    cache: BoundedAllocatorCache | None,
) -> None

set_parallel_size

set_parallel_size(
    parallel_size: int,
    mode: str = "tile",
    process_group: ProcessGroup | None = None,
) -> None