Skip to content

vllm_omni.diffusion.models.magi2.audio_decoder

Native Stable Audio Open decoder used by MAGI-2.

The MAGI checkpoint embeds Stable Audio's Oobleck VAE under the pretransform.model prefix and uses the sequential module names from stable-audio-tools. This module converts those names to Diffusers' AutoencoderOobleck layout and loads only the decoder tensors.

Magi2AudioDecoder

Bases: Module

Decode MAGI-2 audio latents with the bundled Stable Audio VAE weights.

decoder instance-attribute

decoder = autoencoder.decoder.to(
    device=device, dtype=dtype
)

forward class-attribute instance-attribute

forward = decode

latent_fps property

latent_fps: float

sample_rate instance-attribute

sample_rate = int(oobleck_kwargs['sampling_rate'])

decode

decode(latents: Tensor) -> Tensor

convert_stable_audio_decoder_key

convert_stable_audio_decoder_key(key: str) -> str | None

Map one Stable Audio decoder key to Diffusers' Oobleck decoder.

Unrelated checkpoint tensors return None. A key inside the decoder namespace that does not match the released Stable Audio layout raises so a changed checkpoint cannot be loaded partially by accident.

convert_stable_audio_decoder_state_dict

convert_stable_audio_decoder_state_dict(
    state_dict: Mapping[str, Tensor],
) -> dict[str, Tensor]

Convert and filter a Stable Audio checkpoint state dictionary.