Skip to content

vllm_omni.transformers_utils.configs.minimax_music3

HuggingFace config for the MiniMax Music 3 composite repository.

The checkpoint root carries only::

{"architectures": ["MiniMaxMusic3ForConditionalGeneration"],
 "model_type": "minimax_music3"}

and ships no modelling code, so trust_remote_code has nothing to load and AutoConfig rejects the repository outright. Stage 0 sidesteps this by declaring model_subdir="language_model" and reading the plain Qwen3ForCausalLM config, but the acoustic stage must open the root to reach condition_encoder/, transformer/ and vocoder/. Registering this class is what lets that stage build a ModelConfig at all.

The generic transformer fields describe the AR backbone, because the composite root has no transformer of its own: they are metadata for vLLM's config plumbing, not a description of the acoustic stage. The acoustic stage holds no vLLM attention layers, so no KV cache is derived from them; it reads its real geometry from models/minimax_music3/constants.py and its weights from the component folders.

MiniMaxMusic3Config

Bases: PretrainedConfig

Root config for the multi-component MiniMax Music 3 repository.

audio_channels instance-attribute

audio_channels = 2

audio_frame_rate instance-attribute

audio_frame_rate = 25

audio_vocab_size instance-attribute

audio_vocab_size = 1024

c0_vocab_size instance-attribute

c0_vocab_size = 16384

model_type class-attribute instance-attribute

model_type = 'minimax_music3'

num_codebooks instance-attribute

num_codebooks = 8

sample_rate instance-attribute

sample_rate = 32000