Skip to content

vllm_omni.model_executor.models.minimax_music3.pipeline

MiniMax Music 3 pipeline topology: AR talker -> acoustic decoder.

Stage 0 minimax_music3_ar: a Qwen3 backbone (the repo's language_model/ subfolder is a plain Qwen3ForCausalLM) plus a depth decoder, wrapped by a class that owns its sampler. It emits 200-frame windows of frame hidden states and the matching 8-codebook RVQ codes.

Stage 1 minimax_music3_acoustic: flow-matching DiT plus DAV vocoder. Pure compute, no KV cache, weights loaded from the repo's condition_encoder/, transformer/ and vocoder/ subfolders. Output is 32 kHz stereo.

Classifier-free guidance is mandatory for this checkpoint: every request decodes as a cond/uncond pair, and the uncond row is produced by prompt_expand_func. The deploy YAMLs therefore carry default_sampling_params.extra_args with cfg_role and cfg_scale so that has_sampling_extra_args is on for the stage and per-row extra args reach the model's forward.

MINIMAX_MUSIC3_AUDIO_END_TOKEN_ID module-attribute

MINIMAX_MUSIC3_AUDIO_END_TOKEN_ID = 151670

MINIMAX_MUSIC3_PIPELINE module-attribute

MINIMAX_MUSIC3_PIPELINE = PipelineConfig(
    model_type="minimax_music3",
    model_arch="MiniMaxMusic3TalkerForConditionalGeneration",
    hf_architectures=(
        "MiniMaxMusic3ForConditionalGeneration",
    ),
    default_deploy_config_name="minimax_music3.yaml",
    stages=(
        StagePipelineConfig(
            stage_id=0,
            model_stage="minimax_music3_ar",
            execution_type=StageExecutionType.LLM_AR,
            input_sources=(),
            owns_tokenizer=True,
            engine_output_type="latent",
            model_subdir="language_model",
            tokenizer_subdir="tokenizer",
            sampling_constraints={
                "detokenize": False,
                "stop_token_ids": [
                    MINIMAX_MUSIC3_AUDIO_END_TOKEN_ID
                ],
            },
            prompt_expand_func=f"{_PROC}.expand_cfg_prompts",
            async_chunk_process_next_stage_input_func=f"{_PROC}.ar2acoustic_async_chunk",
            extras={
                "tts_args": {
                    "max_instructions_length": _MAX_INSTRUCTIONS_LENGTH
                }
            },
        ),
        StagePipelineConfig(
            stage_id=1,
            model_stage="minimax_music3_acoustic",
            execution_type=StageExecutionType.LLM_GENERATION,
            input_sources=(0,),
            final_output=True,
            final_output_type="audio",
            engine_output_type="audio",
            model_arch="MiniMaxMusic3AcousticForConditionalGeneration",
            tokenizer_subdir="tokenizer",
            retains_state_across_chunks=True,
            sync_process_input_func=f"{_PROC}.ar2acoustic",
            sampling_constraints={"detokenize": True},
        ),
    ),
)