Skip to content

vllm_omni.model_executor.models.minimax_music3.dav

DAC-style vocoder for MiniMax Music 3.

Only the decoder half of the autoencoder ships in the checkpoint, which is all inference needs: the flow-matching stage produces the latent directly, so the encoder and the quantizer never run.

Stereo is folded into the batch. A [B, 128, T] latent is read as two 64-channel mono latents per item, decoded as 2B independent mono streams, and unfolded again at the end. The four transposed-convolution blocks upsample by 8, 8, 4 and 2, so one latent frame becomes 512 samples at 44.1 kHz.

Module names match the vocoder/ component folder key for key, so its state dict loads with strict=True and no remapping. The convolutions keep the checkpoint's weight-norm parameterization (weight_g/weight_v) through loading; :func:remove_weight_norm folds it away afterwards, since the factorization is fixed once the weights stop training.

:class:StreamingResampler carries the 44.1 kHz output down to the 32 kHz response rate across chunk boundaries, which is not something a plain per-chunk resample can do.

logger module-attribute

logger = init_logger(__name__)

DecoderBlock

Bases: Module

One upsampling stage: activate, stretch by stride, then refine.

conv_t1 instance-attribute

conv_t1 = _wn_conv_transpose(
    input_dim,
    output_dim,
    kernel_size=2 * stride,
    stride=stride,
    padding=math.ceil(stride / 2),
)

res_unit1 instance-attribute

res_unit1 = ResidualUnit(output_dim, _RESIDUAL_DILATIONS[0])

res_unit2 instance-attribute

res_unit2 = ResidualUnit(output_dim, _RESIDUAL_DILATIONS[1])

res_unit3 instance-attribute

res_unit3 = ResidualUnit(output_dim, _RESIDUAL_DILATIONS[2])

snake1 instance-attribute

snake1 = Snake1d(input_dim)

forward

forward(x: Tensor) -> Tensor

MiniMaxMusic3Vocoder

Bases: Module

Latent to waveform: [B, 128, T] -> [B, 2, T * 512] at 44.1 kHz.

blocks instance-attribute

blocks = nn.ModuleList(blocks)

conv_in instance-attribute

conv_in = _wn_conv(
    _DECODER_INPUT_DIM,
    _DECODER_HIDDEN_DIM,
    kernel_size=_KERNEL_SIZE,
    padding=_KERNEL_SIZE // 2,
)

conv_out instance-attribute

conv_out = _wn_conv(
    output_dim,
    1,
    kernel_size=_KERNEL_SIZE,
    padding=_KERNEL_SIZE // 2,
)

dec_in_proj instance-attribute

dec_in_proj = nn.Conv1d(
    _STREAM_CHANNELS, _DECODER_INPUT_DIM, kernel_size=1
)

snake_out instance-attribute

snake_out = Snake1d(output_dim)

upsampling_factor property

upsampling_factor: int

Waveform samples produced per latent frame.

forward

forward(latent: Tensor) -> Tensor

Decode a latent to interleaved stereo.

Parameters:

Name Type Description Default
latent Tensor

[B, 128, T] vocoder latent.

required

Returns:

Type Description
Tensor

[B, 2, T * 512] waveform at 44.1 kHz, unclamped.

Raises:

Type Description
ValueError

If the latent is not [B, 128, T].

ResidualUnit

Bases: Module

Dilated snake-gated residual block; preserves length and width.

conv1 instance-attribute

conv1 = _wn_conv(
    dim,
    dim,
    kernel_size=_KERNEL_SIZE,
    dilation=dilation,
    padding=pad,
)

conv2 instance-attribute

conv2 = _wn_conv(dim, dim, kernel_size=1)

snake1 instance-attribute

snake1 = Snake1d(dim)

snake2 instance-attribute

snake2 = Snake1d(dim)

forward

forward(x: Tensor) -> Tensor

Snake1d

Bases: Module

Snake activation with one learned frequency per channel.

alpha instance-attribute

alpha = nn.Parameter(torch.ones(1, channels, 1))

forward

forward(x: Tensor) -> Tensor

StreamingResampler

Resample a chunked stream to the response rate without a seam.

A polyphase resampler is time-invariant but not offset-invariant: it only reproduces the same output grid when the input offset is a whole multiple of orig // gcd(orig, new) samples, and it zero-pads both ends of every call. Resampling each decoded chunk on its own therefore shifts the grid by a fraction of a sample at every join and tapers the signal across it. Measured against a single whole-stream resample on 20 s of broadband noise, that costs a peak error of 2.2 (full scale) and drifts the length.

This buffers instead. Only whole blocks are emitted, so every boundary lands on an exact output sample; the previously emitted tail is prepended as real left context and the not-yet-emitted head appended as real right context, and both are dropped from the result. On the same measurement the emitted stream is bit-identical to the whole-stream resample.

One instance per request. It holds at most a couple of blocks of audio.

channels instance-attribute

channels = int(channels)

new_freq instance-attribute

new_freq = int(new_freq)

orig_freq instance-attribute

orig_freq = int(orig_freq)

push

push(
    chunks: Sequence[Tensor], *, flush: bool = False
) -> Tensor

Append source-rate audio and return whatever is now emittable.

Parameters:

Name Type Description Default
chunks Sequence[Tensor]

New [channels, T] segments at orig_freq, in order.

required
flush bool

Emit the remainder too, because the stream has ended.

False

Returns:

Type Description
Tensor

[channels, T] at new_freq; empty when nothing is ready.

reset

reset() -> None

Drop all buffered audio.

remove_weight_norm

remove_weight_norm(module: Module) -> int

Fold weight normalization into every convolution weight in place.

Returns:

Type Description
int

How many submodules were folded.

snake

snake(x: Tensor, alpha: Tensor) -> Tensor

Periodic x + sin^2(alpha x) / alpha activation.

The reshape to three dimensions is what the reference does and is a no-op for the rank-3 activations used here; it is kept so the numerics are identical for any rank.