Skip to content

vllm_omni.model_executor.models.minimax_music3.weights

Component weight loading for MiniMax Music 3.

The checkpoint is a multi-component repo. The Qwen3 backbone is a normal Qwen3ForCausalLM under language_model/ and loads through vLLM. The audio embedding table and the depth decoder live in rvq_depth_decoder/, and the acoustic stage's three components live in their own folders. None of those are things vLLM's loader knows how to map, so they are read directly.

Every loader here is strict. A silently partial load would produce audio that sounds plausible but is wrong, which is far worse than a startup failure.

CONDITION_ENCODER_DIR module-attribute

CONDITION_ENCODER_DIR = 'condition_encoder'

DEPTH_DECODER_DIR module-attribute

DEPTH_DECODER_DIR = 'rvq_depth_decoder'

TRANSFORMER_DIR module-attribute

TRANSFORMER_DIR = 'transformer'

VOCODER_DIR module-attribute

VOCODER_DIR = 'vocoder'

logger module-attribute

logger = init_logger(__name__)

load_component_state

load_component_state(
    root: Path, component: str, *, device: str = "cpu"
) -> dict[str, Tensor]

Read one component's safetensors, sharded or not.

Raises:

Type Description
FileNotFoundError

If the component has no weight files.

load_depth_decoder_weights

load_depth_decoder_weights(
    vllm_config: VllmConfig,
    *,
    audio_embeddings: Embedding,
    rvq_decoder: Module,
) -> set[str]

Load the audio embedding table and the RVQ depth decoder in place.

Parameters:

Name Type Description Default
vllm_config VllmConfig

The stage's config, used to locate the checkpoint.

required
audio_embeddings Embedding

The c1..c7 embedding table to populate.

required
rvq_decoder Module

The depth decoder to populate.

required

Returns:

Type Description
set[str]

The model-side parameter names that were loaded.

Raises:

Type Description
ValueError

If the audio embedding is absent or the wrong shape.

resolve_repo_root

resolve_repo_root(
    model_path: str | PathLike[str],
    *,
    revision: str | None = None,
    download_dir: str | None = None,
) -> Path

Find the multi-component repository root from a stage's model path.

A stage that declares model_subdir sees the subfolder as its model path, so the root is usually the parent. Both are checked, and the search walks up a couple of levels for snapshot layouts.

The acoustic stage declares no model_subdir because it reads the root itself, so on a hub-id deployment its model path is the repo id rather than a directory. Resolve that to a local snapshot instead of treating it as a relative path.

A cold cache is recoverable too: stage init pre-downloads only language_model/ and tokenizer/, so the talker's model path exists while the component folders do not. Fetch them here instead of failing.

A candidate is accepted only when its components are actually loadable. Matching on folder names alone would accept a snapshot left weightless by an interrupted download and defer the failure to load_component_state, past the point where the repair below could still run.

Raises:

Type Description
FileNotFoundError

If no ancestor looks like the repository root.