vllm_omni.model_executor.models.minimax_music3.dav ¶
DAC-style vocoder for MiniMax Music 3.
Only the decoder half of the autoencoder ships in the checkpoint, which is all inference needs: the flow-matching stage produces the latent directly, so the encoder and the quantizer never run.
Stereo is folded into the batch. A [B, 128, T] latent is read as two 64-channel mono latents per item, decoded as 2B independent mono streams, and unfolded again at the end. The four transposed-convolution blocks upsample by 8, 8, 4 and 2, so one latent frame becomes 512 samples at 44.1 kHz.
Module names match the vocoder/ component folder key for key, so its state dict loads with strict=True and no remapping. The convolutions keep the checkpoint's weight-norm parameterization (weight_g/weight_v) through loading; :func:remove_weight_norm folds it away afterwards, since the factorization is fixed once the weights stop training.
:class:StreamingResampler carries the 44.1 kHz output down to the 32 kHz response rate across chunk boundaries, which is not something a plain per-chunk resample can do.
DecoderBlock ¶
MiniMaxMusic3Vocoder ¶
Bases: Module
Latent to waveform: [B, 128, T] -> [B, 2, T * 512] at 44.1 kHz.
conv_in instance-attribute ¶
conv_in = _wn_conv(
_DECODER_INPUT_DIM,
_DECODER_HIDDEN_DIM,
kernel_size=_KERNEL_SIZE,
padding=_KERNEL_SIZE // 2,
)
conv_out instance-attribute ¶
dec_in_proj instance-attribute ¶
forward ¶
Decode a latent to interleaved stereo.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
latent | Tensor |
| required |
Returns:
| Type | Description |
|---|---|
Tensor |
|
Raises:
| Type | Description |
|---|---|
ValueError | If the latent is not |
ResidualUnit ¶
Bases: Module
Dilated snake-gated residual block; preserves length and width.
conv1 instance-attribute ¶
Snake1d ¶
StreamingResampler ¶
Resample a chunked stream to the response rate without a seam.
A polyphase resampler is time-invariant but not offset-invariant: it only reproduces the same output grid when the input offset is a whole multiple of orig // gcd(orig, new) samples, and it zero-pads both ends of every call. Resampling each decoded chunk on its own therefore shifts the grid by a fraction of a sample at every join and tapers the signal across it. Measured against a single whole-stream resample on 20 s of broadband noise, that costs a peak error of 2.2 (full scale) and drifts the length.
This buffers instead. Only whole blocks are emitted, so every boundary lands on an exact output sample; the previously emitted tail is prepended as real left context and the not-yet-emitted head appended as real right context, and both are dropped from the result. On the same measurement the emitted stream is bit-identical to the whole-stream resample.
One instance per request. It holds at most a couple of blocks of audio.
push ¶
Append source-rate audio and return whatever is now emittable.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
chunks | Sequence[Tensor] | New | required |
flush | bool | Emit the remainder too, because the stream has ended. | False |
Returns:
| Type | Description |
|---|---|
Tensor |
|
remove_weight_norm ¶
remove_weight_norm(module: Module) -> int
Fold weight normalization into every convolution weight in place.
Returns:
| Type | Description |
|---|---|
int | How many submodules were folded. |
snake ¶
Periodic x + sin^2(alpha x) / alpha activation.
The reshape to three dimensions is what the reference does and is a no-op for the rank-3 activations used here; it is kept so the numerics are identical for any rank.