vllm_omni.model_executor.models.minimax_music3.prompt ¶
MiniMax Music 3 prompt construction.
The prompt is a single flat string, not a chat template::
<|im_start|><|caption_start|>{caption}<|caption_end|>
<|lyrics_start|>{lyrics}<|lyrics_end|><|im_end|><|audio_start|>
The classifier-free-guidance twin reuses the same token ids with the whole caption-and-lyrics span (delimiters included) overwritten by <|audio_cfg|>, so only the chat markers and <|audio_start|> survive. Both rows therefore have identical length and identical audio-start position.
SPECIAL_TOKEN_IDS module-attribute ¶
SPECIAL_TOKEN_IDS: dict[str, int] = {
"<|im_start|>": 151644,
"<|im_end|>": 151645,
"<|audio_cfg|>": 151654,
"<|audio_start|>": 151669,
"<|audio_end|>": 151670,
"<|caption_start|>": 151671,
"<|caption_end|>": 151672,
"<|lyrics_start|>": 151673,
"<|lyrics_end|>": 151674,
}
build_cfg_null_token_ids ¶
Return the unconditioned twin's token ids for a conditioned prompt.
Everything between the leading <|im_start|> and the trailing <|im_end|><|audio_start|> becomes <|audio_cfg|>. Length and the audio-start position are preserved, which is what lets the two rows decode in lockstep.
Raises:
| Type | Description |
|---|---|
ValueError | If the prompt is too short to carry a caption and lyrics, or does not have the expected framing tokens. |
clean_caption ¶
Apply the source order: special tokens, then Markdown, then newlines.
<|tag value|> forms are rewritten to tag is value so a caption cannot inject a real special token into the prompt.
normalize_lyrics ¶
Normalize lyrics to the form the checkpoint was trained on.