Skip to content

vllm_omni.model_executor.models.minimax_music3

MiniMax Music 3 text-to-music.

Stage 0 is a Qwen3 backbone with an eight-codebook RVQ frame; stage 1 turns the resulting conditioning frames into 32 kHz stereo audio through a flow-matching transformer and a DAC decoder.

Modules:

Name Description
acoustic

Stage-1 acoustic decoder for MiniMax Music 3.

chunking

Chunk boundary and overlap arithmetic for MiniMax Music 3.

constants

Inference-contract constants for MiniMax Music 3.

dav

DAC-style vocoder for MiniMax Music 3.

depth_graph

CUDA graph capture for the RVQ depth loop.

dit

Condition encoder and flow-matching transformer for MiniMax Music 3.

frame_state

Per-request AR state for MiniMax Music 3.

pipeline

MiniMax Music 3 pipeline topology: AR talker -> acoustic decoder.

prompt

MiniMax Music 3 prompt construction.

rvq_decoder

RVQ depth decoder for MiniMax Music 3 (codebooks c1..c7).

sampling

Deterministic guided sampling for MiniMax Music 3.

staging

Pinned host staging for the small per-step index vectors.

talker

Stage-0 talker for MiniMax Music 3.

weights

Component weight loading for MiniMax Music 3.

AR_CHUNK_FRAMES module-attribute

AR_CHUNK_FRAMES = 200

AR_CHUNK_HOP_FRAMES module-attribute

AR_CHUNK_HOP_FRAMES = 100

AR_HIDDEN_SIZE module-attribute

AUDIO_CODE_OFFSET module-attribute

AUDIO_CODE_OFFSET = 151675

MAX_AUDIO_FRAMES module-attribute

MAX_AUDIO_FRAMES = 9000

OUTPUT_SAMPLE_RATE module-attribute

OUTPUT_SAMPLE_RATE = 32000

SPECIAL_TOKEN_IDS module-attribute

SPECIAL_TOKEN_IDS: dict[str, int] = {
    "<|im_start|>": 151644,
    "<|im_end|>": 151645,
    "<|audio_cfg|>": 151654,
    "<|audio_start|>": 151669,
    "<|audio_end|>": 151670,
    "<|caption_start|>": 151671,
    "<|caption_end|>": 151672,
    "<|lyrics_start|>": 151673,
    "<|lyrics_end|>": 151674,
}

build_cfg_null_token_ids

build_cfg_null_token_ids(
    prompt_token_ids: list[int],
) -> list[int]

Return the unconditioned twin's token ids for a conditioned prompt.

Everything between the leading <|im_start|> and the trailing <|im_end|><|audio_start|> becomes <|audio_cfg|>. Length and the audio-start position are preserved, which is what lets the two rows decode in lockstep.

Raises:

Type Description
ValueError

If the prompt is too short to carry a caption and lyrics, or does not have the expected framing tokens.

build_prompt

build_prompt(caption: str, lyrics: str) -> str

Build the conditioned prompt string.