vllm_omni.model_executor.models.personaplex.duplex ¶
PersonaPlex full-duplex integration.
PersonaPlex (nvidia/personaplex-7b-v1) is a Moshi finetune: a pure-lockstep speech-to-speech model. This package plugs it into the generic duplex serving stack through the standard plugin seams (duplex_serving_adapter / duplex_runtime_extension dotted strings in the model's pipeline.py):
- :class:
PersonaPlexConfigimmutable session config (voice / persona / sampling) - :class:
PersonaPlexServingRuntimeAdaptertheServingRuntimeAdapterimpl - :class:
PersonaPlexDuplexRuntimeExtensionthe engineDuplexRuntimeExtension - :class:
PersonaPlexStage0DuplexRuntimeStage 0 session state and prefill - :class:
PersonaPlexPcmAppendBufferPCM input framing
Modules:
| Name | Description |
|---|---|
config | Configuration for the PersonaPlex full-duplex backend. |
data_plane | |
input | |
policy | PersonaPlex frame/token contract (the model policy, no engine state). |
runtime_extension | |
serving_adapter | |
stage0 | |
PersonaPlexConfig dataclass ¶
Immutable session configuration for a PersonaPlex conversation.
Attributes:
| Name | Type | Description |
|---|---|---|
hf_repo | str | HuggingFace repo holding the weights, Mimi codec and tokenizer. |
voice_prompt | str | Voice-clone reference. Either a bundled basename ( |
persona | str | System role text; injected as |
device | str | Torch device for the backend ( |
cpu_offload | bool | Offload LM layers to CPU when GPU memory is tight (needs |
batch_size | int | Concurrent conversation slots sharing one engine. |
Note: the native stepper decodes greedily (argmax) for both the text head and the depformer, so there are no sampling knobs here yet. Temperature / top-k / seed fields will be added if and when a sampling path is wired in.
PersonaPlexDuplexRuntimeExtension ¶
PersonaPlex policy for the engine-owned resumable duplex request.
configure_sampling_params ¶
configure_sampling_params(
*,
runtime_config: dict[str, Any],
defaults: tuple[object, ...],
) -> tuple[object, ...]
PersonaPlexPcmAppendBuffer ¶
Transactionally frame 24 kHz float PCM into PersonaPlex 80 ms units.
PersonaPlexServingRuntimeAdapter ¶
private_runtime_config_keys class-attribute instance-attribute ¶
session_states instance-attribute ¶
session_states: dict[
str, PersonaPlexServingSessionState
] = {}
data_plane_context staticmethod ¶
data_plane_context(
*,
epoch: int,
turn_id: int,
active_response_turn_id: int | None,
active_response_id: str | None,
auto_responds: bool,
response_format: str,
speed: float | None,
modalities: tuple[str, ...],
) -> PersonaPlexDataPlaneContext
prepare_runtime_config async classmethod ¶
runtime_config_for_update classmethod ¶
PersonaPlexStage0DuplexRuntime ¶
PrefillStep dataclass ¶
One tick of a recycled slot's system-prompt replay (see batched serving).