vllm_omni.model_executor.models.minimax_music3.acoustic ¶
Stage-1 acoustic decoder for MiniMax Music 3.
Consumes the AR stage's conditioning spans and returns 32 kHz stereo audio. Each span is cut into 200-frame windows that overlap by half; every window is solved by the flow-matching DiT with the previous window's tail latent pinned as a prompt, decoded by the vocoder, and cropped so that concatenating the survivors reconstructs the song exactly once.
The window arithmetic lives in :mod:.chunking and is driven entirely by the span's global frame range, so the streaming path (one window per engine step) and the non-streaming path (the whole song in one payload) produce identical audio for the same seed.
The stage runs in float32. TF32 keeps that affordable, and bfloat16 measurably degrades a 30-step solve.
MiniMaxMusic3AcousticForConditionalGeneration ¶
Bases: Module
Flow-matching DiT plus DAC vocoder, driven by the generation runner.
dit_cfg_scale instance-attribute ¶
dit_cfg_scale = (
float(cfg_scale)
if isinstance(cfg_scale, int | float)
and not isinstance(cfg_scale, bool)
else DEFAULT_DIT_CFG_SCALE
)
dit_steps instance-attribute ¶
dit_steps = max(
1, _meta_int(extra.get("dit_steps"), DEFAULT_DIT_STEPS)
)
enable_update_additional_information instance-attribute ¶
compute_logits ¶
compute_logits(
hidden_states: Tensor | OmniOutput,
sampling_metadata: Any = None,
) -> None
embed_input_ids ¶
embed_input_ids(input_ids: Tensor, **_: Any) -> Tensor
Stable dummy embedding: this stage never reads token embeddings.
forward ¶
forward(
input_ids: Tensor | None = None,
positions: Tensor | None = None,
intermediate_tensors: Any = None,
inputs_embeds: Tensor | None = None,
runtime_additional_information: list[dict[str, Any]]
| None = None,
**kwargs: Any,
) -> OmniOutput
Decode this step's spans into one stereo waveform per request.
Raises:
| Type | Description |
|---|---|
ValueError | If a payload is malformed. Failing the step is deliberate: a request whose conditioning cannot be decoded has no audio to return, and reporting an error is better than leaving the caller waiting on a stream that will never end. |
load_weights ¶
Load the three acoustic components from their own folders.
The iterator vLLM hands over covers the stage's own checkpoint folder, which holds nothing this stage needs. It is drained anyway: the loader streams from a file handle the caller closes only once the iterator is exhausted, so abandoning it wedges startup.
make_omni_output ¶
make_omni_output(
model_outputs: Tensor | OmniOutput | tuple,
**kwargs: Any,
) -> OmniOutput