vllm_omni.model_executor.models.minimax_music3.pipeline ¶
MiniMax Music 3 pipeline topology: AR talker -> acoustic decoder.
Stage 0 minimax_music3_ar: a Qwen3 backbone (the repo's language_model/ subfolder is a plain Qwen3ForCausalLM) plus a depth decoder, wrapped by a class that owns its sampler. It emits 200-frame windows of frame hidden states and the matching 8-codebook RVQ codes.
Stage 1 minimax_music3_acoustic: flow-matching DiT plus DAV vocoder. Pure compute, no KV cache, weights loaded from the repo's condition_encoder/, transformer/ and vocoder/ subfolders. Output is 32 kHz stereo.
Classifier-free guidance is mandatory for this checkpoint: every request decodes as a cond/uncond pair, and the uncond row is produced by prompt_expand_func. The deploy YAMLs therefore carry default_sampling_params.extra_args with cfg_role and cfg_scale so that has_sampling_extra_args is on for the stage and per-row extra args reach the model's forward.
MINIMAX_MUSIC3_PIPELINE module-attribute ¶
MINIMAX_MUSIC3_PIPELINE = PipelineConfig(
model_type="minimax_music3",
model_arch="MiniMaxMusic3TalkerForConditionalGeneration",
hf_architectures=(
"MiniMaxMusic3ForConditionalGeneration",
),
default_deploy_config_name="minimax_music3.yaml",
stages=(
StagePipelineConfig(
stage_id=0,
model_stage="minimax_music3_ar",
execution_type=StageExecutionType.LLM_AR,
input_sources=(),
owns_tokenizer=True,
engine_output_type="latent",
model_subdir="language_model",
tokenizer_subdir="tokenizer",
sampling_constraints={
"detokenize": False,
"stop_token_ids": [
MINIMAX_MUSIC3_AUDIO_END_TOKEN_ID
],
},
prompt_expand_func=f"{_PROC}.expand_cfg_prompts",
async_chunk_process_next_stage_input_func=f"{_PROC}.ar2acoustic_async_chunk",
extras={
"tts_args": {
"max_instructions_length": _MAX_INSTRUCTIONS_LENGTH
}
},
),
StagePipelineConfig(
stage_id=1,
model_stage="minimax_music3_acoustic",
execution_type=StageExecutionType.LLM_GENERATION,
input_sources=(0,),
final_output=True,
final_output_type="audio",
engine_output_type="audio",
model_arch="MiniMaxMusic3AcousticForConditionalGeneration",
tokenizer_subdir="tokenizer",
retains_state_across_chunks=True,
sync_process_input_func=f"{_PROC}.ar2acoustic",
sampling_constraints={"detokenize": True},
),
),
)