Skip to content

vllm_omni.model_executor.models.minimax_music3.acoustic

Stage-1 acoustic decoder for MiniMax Music 3.

Consumes the AR stage's conditioning spans and returns 32 kHz stereo audio. Each span is cut into 200-frame windows that overlap by half; every window is solved by the flow-matching DiT with the previous window's tail latent pinned as a prompt, decoded by the vocoder, and cropped so that concatenating the survivors reconstructs the song exactly once.

The window arithmetic lives in :mod:.chunking and is driven entirely by the span's global frame range, so the streaming path (one window per engine step) and the non-streaming path (the whole song in one payload) produce identical audio for the same seed.

The stage runs in float32. TF32 keeps that affordable, and bfloat16 measurably degrades a 30-step solve.

logger module-attribute

logger = init_logger(__name__)

MiniMaxMusic3AcousticForConditionalGeneration

Bases: Module

Flow-matching DiT plus DAC vocoder, driven by the generation runner.

dit instance-attribute

dit_cfg_scale instance-attribute

dit_cfg_scale = (
    float(cfg_scale)
    if isinstance(cfg_scale, int | float)
    and not isinstance(cfg_scale, bool)
    else DEFAULT_DIT_CFG_SCALE
)

dit_steps instance-attribute

dit_steps = max(
    1, _meta_int(extra.get("dit_steps"), DEFAULT_DIT_STEPS)
)

enable_update_additional_information instance-attribute

enable_update_additional_information = True

has_postprocess instance-attribute

has_postprocess = False

has_preprocess instance-attribute

has_preprocess = False

have_multimodal_outputs instance-attribute

have_multimodal_outputs = True

input_modalities class-attribute instance-attribute

input_modalities = 'audio'

model_path instance-attribute

model_path = vllm_config.model_config.model

requires_raw_input_tokens instance-attribute

requires_raw_input_tokens = True

vllm_config instance-attribute

vllm_config = vllm_config

vocoder instance-attribute

compute_logits

compute_logits(
    hidden_states: Tensor | OmniOutput,
    sampling_metadata: Any = None,
) -> None

embed_input_ids

embed_input_ids(input_ids: Tensor, **_: Any) -> Tensor

Stable dummy embedding: this stage never reads token embeddings.

forward

forward(
    input_ids: Tensor | None = None,
    positions: Tensor | None = None,
    intermediate_tensors: Any = None,
    inputs_embeds: Tensor | None = None,
    runtime_additional_information: list[dict[str, Any]]
    | None = None,
    **kwargs: Any,
) -> OmniOutput

Decode this step's spans into one stereo waveform per request.

Raises:

Type Description
ValueError

If a payload is malformed. Failing the step is deliberate: a request whose conditioning cannot be decoded has no audio to return, and reporting an error is better than leaving the caller waiting on a stream that will never end.

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

Load the three acoustic components from their own folders.

The iterator vLLM hands over covers the stage's own checkpoint folder, which holds nothing this stage needs. It is drained anyway: the loader streams from a file handle the caller closes only once the iterator is exhausted, so abandoning it wedges startup.

make_omni_output

make_omni_output(
    model_outputs: Tensor | OmniOutput | tuple,
    **kwargs: Any,
) -> OmniOutput

on_requests_finished

on_requests_finished(
    finished_req_ids: Iterable[str],
) -> None

Drop the streaming state of finished or aborted requests.