Skip to content

vllm_omni.model_executor.models.minimax_music3.sampling

Deterministic guided sampling for MiniMax Music 3.

Two things happen per code, in this order:

  1. Classifier-free guidance. Every request decodes twice, once on the real prompt and once on a prompt whose caption and lyrics have been replaced by <|audio_cfg|>. Codes are drawn from uncond + (cond - uncond) * 1.5. For c0 the candidate set is additionally masked to the conditioned row's own top 50, so the guided draw cannot wander outside what the conditioned branch already considered plausible.
  2. Seeded top-k. A batched, order-independent categorical draw: each row gets its own RNG stream keyed by (seed, position), so a request's audio depends only on its own seed and never on who else happens to share the batch.

The RNG is a MurmurHash3 x86_32 of (seed_lo, seed_hi, position, column) turned into Gumbel noise, and the draw is an argmax over logits + gumbel. That is mathematically a categorical sample from softmax(logits), but unlike torch.multinomial it needs no per-row generator state, so it is CUDA-graph safe and reproducible across batch compositions.

apply_c0_cfg

apply_c0_cfg(
    cond: Tensor,
    uncond: Tensor,
    *,
    scale: float = AR_CFG_SCALE,
) -> Tensor

Guide cond by uncond, restricted to cond's own top candidates.

The mask is a second, separate restriction from the top-k the sampler applies afterwards. Dropping it widens the distribution the model draws from and audibly loosens how closely the output follows the caption.

apply_depth_cfg

apply_depth_cfg(
    cond: Tensor,
    uncond: Tensor,
    *,
    scale: float = AR_CFG_SCALE,
) -> Tensor

Guide the depth heads. Unlike c0 these are not additionally masked.

draw_from_gumbel

draw_from_gumbel(
    logits: Tensor, gumbel: Tensor, *, top_k: int = AR_TOP_K
) -> Tensor

Argmax of top_k-masked logits plus precomputed Gumbel noise.

Parameters:

Name Type Description Default
logits Tensor

[rows, vocab] unnormalized scores.

required
gumbel Tensor

[rows, vocab] float64 noise from :func:gumbel_noise.

required
top_k int

Candidates retained per row before the draw.

AR_TOP_K

Returns:

Type Description
Tensor

[rows] sampled column indices.

gumbel_noise

gumbel_noise(
    seeds: Tensor, positions: Tensor, num_columns: int
) -> Tensor

Return the Gumbel noise for one draw per position, [..., cols].

Parameters:

Name Type Description Default
seeds Tensor

[rows] per-request seeds.

required
positions Tensor

[rows] or [rows, draws] draw identifiers.

required
num_columns int

Candidate count, i.e. the width of the logits.

required

Returns:

Type Description
Tensor

float64 noise shaped positions.shape + (num_columns,).

murmur_hash32

murmur_hash32(
    seeds: Tensor, positions: Tensor, columns: Tensor
) -> Tensor

Hash (seed, position, column) to a uint32 held in int64.

The 64-bit seed is consumed as two 32-bit blocks, then the position and the column, for a hashed length of 16 bytes.

Parameters:

Name Type Description Default
seeds Tensor

[rows] per-request seeds.

required
positions Tensor

[rows] draw identifiers, unique per (frame, codebook).

required
columns Tensor

[cols] candidate indices, which must be arange(cols).

required

Returns:

Type Description
Tensor

[rows, cols] int64 tensor holding uint32 values.

row_hash_state

row_hash_state(seeds: Tensor, positions: Tensor) -> Tensor

Hash state after the seed and position blocks, before the columns.

Parameters:

Name Type Description Default
seeds Tensor

[rows] per-request seeds.

required
positions Tensor

[rows] or [rows, draws] draw identifiers. The extra dimension lets every codebook of a frame be hashed in one batch instead of one dispatch per codebook; the arithmetic is elementwise integer, so a batched call is bit-identical to the loop it replaces.

required

Returns:

Type Description
Tensor

A tensor shaped like positions, int64 holding uint32.

sample_topk_seeded

sample_topk_seeded(
    logits: Tensor,
    seeds: Tensor,
    positions: Tensor,
    *,
    top_k: int = AR_TOP_K,
) -> Tensor

Draw one column per row from its own top-k, under its own RNG stream.

Parameters:

Name Type Description Default
logits Tensor

[rows, vocab] unnormalized scores.

required
seeds Tensor

[rows] per-request sampling seeds.

required
positions Tensor

[rows] draw identifiers, unique per (frame, codebook).

required
top_k int

Candidates retained per row before the draw.

AR_TOP_K

Returns:

Type Description
Tensor

[rows] sampled column indices.