vllm_omni.model_executor.models.minimax_music3 ¶
MiniMax Music 3 text-to-music.
Stage 0 is a Qwen3 backbone with an eight-codebook RVQ frame; stage 1 turns the resulting conditioning frames into 32 kHz stereo audio through a flow-matching transformer and a DAC decoder.
Modules:
| Name | Description |
|---|---|
acoustic | Stage-1 acoustic decoder for MiniMax Music 3. |
chunking | Chunk boundary and overlap arithmetic for MiniMax Music 3. |
constants | Inference-contract constants for MiniMax Music 3. |
dav | DAC-style vocoder for MiniMax Music 3. |
depth_graph | CUDA graph capture for the RVQ depth loop. |
dit | Condition encoder and flow-matching transformer for MiniMax Music 3. |
frame_state | Per-request AR state for MiniMax Music 3. |
pipeline | MiniMax Music 3 pipeline topology: AR talker -> acoustic decoder. |
prompt | MiniMax Music 3 prompt construction. |
rvq_decoder | RVQ depth decoder for MiniMax Music 3 (codebooks c1..c7). |
sampling | Deterministic guided sampling for MiniMax Music 3. |
staging | Pinned host staging for the small per-step index vectors. |
talker | Stage-0 talker for MiniMax Music 3. |
weights | Component weight loading for MiniMax Music 3. |
SPECIAL_TOKEN_IDS module-attribute ¶
SPECIAL_TOKEN_IDS: dict[str, int] = {
"<|im_start|>": 151644,
"<|im_end|>": 151645,
"<|audio_cfg|>": 151654,
"<|audio_start|>": 151669,
"<|audio_end|>": 151670,
"<|caption_start|>": 151671,
"<|caption_end|>": 151672,
"<|lyrics_start|>": 151673,
"<|lyrics_end|>": 151674,
}
build_cfg_null_token_ids ¶
Return the unconditioned twin's token ids for a conditioned prompt.
Everything between the leading <|im_start|> and the trailing <|im_end|><|audio_start|> becomes <|audio_cfg|>. Length and the audio-start position are preserved, which is what lets the two rows decode in lockstep.
Raises:
| Type | Description |
|---|---|
ValueError | If the prompt is too short to carry a caption and lyrics, or does not have the expected framing tokens. |