vllm_omni.model_executor.models.personaplex.duplex.config ¶
Configuration for the PersonaPlex full-duplex backend.
PersonaPlex (nvidia/personaplex-7b-v1) is a Moshi finetune: a pure-lockstep full-duplex speech-to-speech model running at the Mimi codec frame rate (12.5 Hz / 80 ms). One config drives both the offline driver and the duplex adapter. Defaults mirror the PersonaPlex reference loop.
DEFAULT_PERSONA module-attribute ¶
DEFAULT_PERSONA = "You are a wise and friendly teacher. Answer questions or provide advice in a clear and engaging way."
PersonaPlexConfig dataclass ¶
Immutable session configuration for a PersonaPlex conversation.
Attributes:
| Name | Type | Description |
|---|---|---|
hf_repo | str | HuggingFace repo holding the weights, Mimi codec and tokenizer. |
voice_prompt | str | Voice-clone reference. Either a bundled basename ( |
persona | str | System role text; injected as |
device | str | Torch device for the backend ( |
cpu_offload | bool | Offload LM layers to CPU when GPU memory is tight (needs |
batch_size | int | Concurrent conversation slots sharing one engine. |
Note: the native stepper decodes greedily (argmax) for both the text head and the depformer, so there are no sampling knobs here yet. Temperature / top-k / seed fields will be added if and when a sampling path is wired in.