vllm_omni.model_executor.models.personaplex.configuration_personaplex ¶
Configuration for PersonaPlex (a Moshi finetune; 2-stage audio->audio pipeline).
PersonaPlex is a staged AR speech model composed of:
- a temporal transformer (a Llama-variant Helium backbone) that predicts, per frame, the next text token and the codebook-0 audio token;
- a depformer (a small per-step transformer) that, conditioned on the temporal hidden state, predicts the remaining audio codebooks;
- the Mimi neural audio codec, which turns the audio codebooks into 24 kHz PCM.
The temporal-transformer hyperparameters live in :class:~vllm_omni.model_executor.models.personaplex.configuration_helium.HeliumConfig (already used by the talker the lead is building). This module reuses it as the temporal_config sub-config so the config tree carries a single source of truth, and adds two further sub-configs (depformer_config and mimi_config).
All defaults are measured from the PersonaPlex checkpoint, whose own config.json is empty; do not treat the Moshi higher-level kwargs as authoritative.
HeliumConfig ¶
Bases: PretrainedConfig
Minimal HF config for the Moshi temporal LM backbone.
The defaults are measured from the PersonaPlex checkpoint rather than inferred from Moshi's higher-level kwargs.
keys_to_ignore_at_inference class-attribute instance-attribute ¶
PersonaPlexConfig ¶
Bases: PretrainedConfig
Top-level configuration for PersonaPlexTalkerForConditionalGeneration.
Mirrors the Qwen3-TTS config layout: a top-level config holding sub-configs for each component. The temporal transformer config is the text config that vLLM consumes (it exposes hidden_size / num_attention_heads and drives the talker's vLLM Llama backbone).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
temporal_config | `dict` or `HeliumConfig`, *optional* | The temporal-transformer (Helium) backbone config. | None |
depformer_config | `dict` or `PersonaPlexDepformerConfig`, *optional* | The depformer (per-step code predictor) config. | None |
mimi_config | `dict` or `PersonaPlexMimiConfig`, *optional* | The Mimi codec config. | None |
text_vocab_size | `int`, *optional*, defaults to 32000 | Text / | 32000 |
text_embedding_rows | `int`, *optional*, defaults to 32001 | Number of rows in the text embedding table (one extra padding row). | 32001 |
audio_vocab_size | `int`, *optional*, defaults to 2048 | Per-codebook audio cardinality ( | 2048 |
num_audio_codebooks | `int`, *optional*, defaults to 16 | Total number of audio codebooks ( | 16 |
mimi_name | `str`, *optional* | Convenience mirror of | None |
depformer_config instance-attribute ¶
depformer_config = self._coerce(
depformer_config, PersonaPlexDepformerConfig
)
sample_rate property ¶
sample_rate: int
Output PCM sample rate in Hz (delegates to the Mimi config).
sub_configs class-attribute instance-attribute ¶
sub_configs = {
"temporal_config": HeliumConfig,
"depformer_config": PersonaPlexDepformerConfig,
"mimi_config": PersonaPlexMimiConfig,
}
temporal_config instance-attribute ¶
temporal_config = self._coerce(
temporal_config, HeliumConfig
)
PersonaPlexDepformerConfig ¶
Bases: PretrainedConfig
Configuration for the PersonaPlex depformer (per-step code predictor).
The depformer is a small autoregressive transformer that runs dep_q inner steps per temporal frame, conditioned on the temporal hidden state, to predict the audio codebooks. Only the first num_active_codebooks codebooks (cb 0..7) are decoded to PCM by Mimi; the remaining codebooks up to dep_q are predicted but not vocoded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
hidden_size | `int`, *optional*, defaults to 1024 | Dimension of the depformer hidden representations. | 1024 |
num_hidden_layers | `int`, *optional*, defaults to 6 | Number of depformer transformer layers. | 6 |
num_attention_heads | `int`, *optional*, defaults to 16 | Number of attention heads per depformer layer. | 16 |
head_dim | `int`, *optional*, defaults to 64 | Per-head attention dimension ( | 64 |
max_position_embeddings | `int`, *optional*, defaults to 8 | The depformer context length ( | 8 |
dep_q | `int`, *optional*, defaults to 16 | Number of audio codebooks the depformer predicts per frame. | 16 |
num_active_codebooks | `int`, *optional*, defaults to 8 | Number of leading codebooks actually decoded to PCM by Mimi. | 8 |
card | `int`, *optional*, defaults to 2048 | Per-codebook cardinality (audio vocab size). | 2048 |
rope_theta | `float`, *optional*, defaults to 10000.0 | The base period of the RoPE embeddings. | 10000.0 |
rms_norm_eps | `float`, *optional*, defaults to 1e-8 | The epsilon used by the fp32 RMS normalization layers. | 1e-08 |
hidden_act | `str`, *optional*, defaults to `"silu"` | SwiGLU gate activation. | 'silu' |
attention_bias | `bool`, *optional*, defaults to `False` | Whether attention projections carry a bias. | False |
mlp_bias | `bool`, *optional*, defaults to `False` | Whether the MLP projections carry a bias. | False |
keys_to_ignore_at_inference class-attribute instance-attribute ¶
PersonaPlexMimiConfig ¶
Bases: PretrainedConfig
Configuration for the Mimi neural audio codec used by PersonaPlex.
The Mimi weights and module live in the external moshi package; this config only carries the scalars the vllm-omni serving layer needs to size buffers and compute audio durations. The actual decoder is instantiated by :class:PersonaPlexCode2Wav via moshi.models.loaders.get_mimi.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_rate | `int`, *optional*, defaults to 24000 | Output PCM sample rate in Hz. | 24000 |
frame_rate | `float`, *optional*, defaults to 12.5 | Mimi codec frame rate in Hz (one frame every 80 ms). | 12.5 |
samples_per_frame | `int`, *optional*, defaults to 1920 | PCM samples produced per codec frame ( | 1920 |
num_codebooks | `int`, *optional*, defaults to 8 | Number of active audio codebooks decoded to PCM ( | 8 |
card | `int`, *optional*, defaults to 2048 | Per-codebook cardinality. | 2048 |
num_channels | `int`, *optional*, defaults to 1 | Number of output audio channels (mono). | 1 |
mimi_name | `str`, *optional* | Filename of the Mimi weight checkpoint inside the model repo. When | None |