vllm_omni.model_executor.models.personaplex ¶
vllm-omni integration for PersonaPlex (a Moshi finetune, full-duplex S2S).
Offline/batch runs go through a two-stage audio->audio pipeline (talker -> code2wav); real-time conversation is served over the duplex WebSocket path.
Modules:
| Name | Description |
|---|---|
configuration_helium | Configuration for the PersonaPlex Helium temporal transformer. |
configuration_personaplex | Configuration for PersonaPlex (a Moshi finetune; 2-stage audio->audio pipeline). |
duplex | PersonaPlex full-duplex integration. |
modeling_helium | vLLM-native PersonaPlex Helium temporal transformer. |
personaplex_code2wav | Stage-1 Code2Wav model for PersonaPlex (Moshi finetune). |
personaplex_depformer | PersonaPlex depformer: the per-step audio code predictor. |
personaplex_embeddings | PersonaPlex input embeddings (Moshi |
personaplex_mimi | Streaming Mimi codec for PersonaPlex (moshi-free). |
personaplex_talker | PersonaPlex talker: the temporal transformer as a vLLM-native omni AR stage. |
personaplex_temporal | Streaming Helium temporal transformer for PersonaPlex (moshi-free, plain torch). |
pipeline | PersonaPlex pipeline: Talker (AR decode -> Mimi codebooks) -> Code2Wav (codebooks -> 24 kHz PCM). |
PersonaPlexConfig ¶
Bases: PretrainedConfig
Top-level configuration for PersonaPlexTalkerForConditionalGeneration.
Mirrors the Qwen3-TTS config layout: a top-level config holding sub-configs for each component. The temporal transformer config is the text config that vLLM consumes (it exposes hidden_size / num_attention_heads and drives the talker's vLLM Llama backbone).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
temporal_config | `dict` or `HeliumConfig`, *optional* | The temporal-transformer (Helium) backbone config. | None |
depformer_config | `dict` or `PersonaPlexDepformerConfig`, *optional* | The depformer (per-step code predictor) config. | None |
mimi_config | `dict` or `PersonaPlexMimiConfig`, *optional* | The Mimi codec config. | None |
text_vocab_size | `int`, *optional*, defaults to 32000 | Text / | 32000 |
text_embedding_rows | `int`, *optional*, defaults to 32001 | Number of rows in the text embedding table (one extra padding row). | 32001 |
audio_vocab_size | `int`, *optional*, defaults to 2048 | Per-codebook audio cardinality ( | 2048 |
num_audio_codebooks | `int`, *optional*, defaults to 16 | Total number of audio codebooks ( | 16 |
mimi_name | `str`, *optional* | Convenience mirror of | None |
depformer_config instance-attribute ¶
depformer_config = self._coerce(
depformer_config, PersonaPlexDepformerConfig
)
sample_rate property ¶
sample_rate: int
Output PCM sample rate in Hz (delegates to the Mimi config).
sub_configs class-attribute instance-attribute ¶
sub_configs = {
"temporal_config": HeliumConfig,
"depformer_config": PersonaPlexDepformerConfig,
"mimi_config": PersonaPlexMimiConfig,
}
temporal_config instance-attribute ¶
temporal_config = self._coerce(
temporal_config, HeliumConfig
)
PersonaPlexDepformerConfig ¶
Bases: PretrainedConfig
Configuration for the PersonaPlex depformer (per-step code predictor).
The depformer is a small autoregressive transformer that runs dep_q inner steps per temporal frame, conditioned on the temporal hidden state, to predict the audio codebooks. Only the first num_active_codebooks codebooks (cb 0..7) are decoded to PCM by Mimi; the remaining codebooks up to dep_q are predicted but not vocoded.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
hidden_size | `int`, *optional*, defaults to 1024 | Dimension of the depformer hidden representations. | 1024 |
num_hidden_layers | `int`, *optional*, defaults to 6 | Number of depformer transformer layers. | 6 |
num_attention_heads | `int`, *optional*, defaults to 16 | Number of attention heads per depformer layer. | 16 |
head_dim | `int`, *optional*, defaults to 64 | Per-head attention dimension ( | 64 |
max_position_embeddings | `int`, *optional*, defaults to 8 | The depformer context length ( | 8 |
dep_q | `int`, *optional*, defaults to 16 | Number of audio codebooks the depformer predicts per frame. | 16 |
num_active_codebooks | `int`, *optional*, defaults to 8 | Number of leading codebooks actually decoded to PCM by Mimi. | 8 |
card | `int`, *optional*, defaults to 2048 | Per-codebook cardinality (audio vocab size). | 2048 |
rope_theta | `float`, *optional*, defaults to 10000.0 | The base period of the RoPE embeddings. | 10000.0 |
rms_norm_eps | `float`, *optional*, defaults to 1e-8 | The epsilon used by the fp32 RMS normalization layers. | 1e-08 |
hidden_act | `str`, *optional*, defaults to `"silu"` | SwiGLU gate activation. | 'silu' |
attention_bias | `bool`, *optional*, defaults to `False` | Whether attention projections carry a bias. | False |
mlp_bias | `bool`, *optional*, defaults to `False` | Whether the MLP projections carry a bias. | False |
keys_to_ignore_at_inference class-attribute instance-attribute ¶
PersonaPlexMimiConfig ¶
Bases: PretrainedConfig
Configuration for the Mimi neural audio codec used by PersonaPlex.
The Mimi weights and module live in the external moshi package; this config only carries the scalars the vllm-omni serving layer needs to size buffers and compute audio durations. The actual decoder is instantiated by :class:PersonaPlexCode2Wav via moshi.models.loaders.get_mimi.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_rate | `int`, *optional*, defaults to 24000 | Output PCM sample rate in Hz. | 24000 |
frame_rate | `float`, *optional*, defaults to 12.5 | Mimi codec frame rate in Hz (one frame every 80 ms). | 12.5 |
samples_per_frame | `int`, *optional*, defaults to 1920 | PCM samples produced per codec frame ( | 1920 |
num_codebooks | `int`, *optional*, defaults to 8 | Number of active audio codebooks decoded to PCM ( | 8 |
card | `int`, *optional*, defaults to 2048 | Per-codebook cardinality. | 2048 |
num_channels | `int`, *optional*, defaults to 1 | Number of output audio channels (mono). | 1 |
mimi_name | `str`, *optional* | Filename of the Mimi weight checkpoint inside the model repo. When | None |