Skip to content

vllm_omni.model_executor.models.personaplex.duplex

PersonaPlex full-duplex integration.

PersonaPlex (nvidia/personaplex-7b-v1) is a Moshi finetune: a pure-lockstep speech-to-speech model. This package plugs it into the generic duplex serving stack through the standard plugin seams (duplex_serving_adapter / duplex_runtime_extension dotted strings in the model's pipeline.py):

  • :class:PersonaPlexConfig immutable session config (voice / persona / sampling)
  • :class:PersonaPlexServingRuntimeAdapter the ServingRuntimeAdapter impl
  • :class:PersonaPlexDuplexRuntimeExtension the engine DuplexRuntimeExtension
  • :class:PersonaPlexStage0DuplexRuntime Stage 0 session state and prefill
  • :class:PersonaPlexPcmAppendBuffer PCM input framing

Modules:

Name Description
config

Configuration for the PersonaPlex full-duplex backend.

data_plane
input
policy

PersonaPlex frame/token contract (the model policy, no engine state).

runtime_extension
serving_adapter
stage0

PersonaPlexConfig dataclass

Immutable session configuration for a PersonaPlex conversation.

Attributes:

Name Type Description
hf_repo str

HuggingFace repo holding the weights, Mimi codec and tokenizer.

voice_prompt str

Voice-clone reference. Either a bundled basename ("NATF2.pt" / "NATM1.pt" from voices.tgz) or a path to a .pt embedding bundle or a reference .wav.

persona str

System role text; injected as <system> ... <system> into the inner-monologue stream at session start.

device str

Torch device for the backend ("cuda" / "cpu").

cpu_offload bool

Offload LM layers to CPU when GPU memory is tight (needs accelerate).

batch_size int

Concurrent conversation slots sharing one engine. 1 is the single-session path; > 1 enables elastic batching with per-slot recycle for new callers.

Note: the native stepper decodes greedily (argmax) for both the text head and the depformer, so there are no sampling knobs here yet. Temperature / top-k / seed fields will be added if and when a sampling path is wired in.

batch_size class-attribute instance-attribute

batch_size: int = 1

cpu_offload class-attribute instance-attribute

cpu_offload: bool = False

device class-attribute instance-attribute

device: str = 'cuda'

frame_size property

frame_size: int

hf_repo class-attribute instance-attribute

hf_repo: str = 'nvidia/personaplex-7b-v1'

persona class-attribute instance-attribute

persona: str = DEFAULT_PERSONA

sample_rate property

sample_rate: int

voice_prompt class-attribute instance-attribute

voice_prompt: str = 'NATF2.pt'

PersonaPlexDuplexRuntimeExtension

PersonaPlex policy for the engine-owned resumable duplex request.

configure_sampling_params

configure_sampling_params(
    *,
    runtime_config: dict[str, Any],
    defaults: tuple[object, ...],
) -> tuple[object, ...]

decide_output

decide_output(
    *,
    stage_id: int,
    final_stage_id: int,
    segment_finished: bool,
    segment_token_ids: tuple[int, ...],
    segment_output_metadata: dict[str, Any],
    output: object,
) -> DuplexOutputDecision | None

plan_append

plan_append(
    *,
    request_id: str,
    fence: DuplexFence,
    session_config: dict[str, Any],
    runtime_config: dict[str, Any],
    seq: int,
    turn_seq: int,
    mode: DuplexInputMode,
    payload: object,
    final: bool,
    sampling_params: object,
) -> DuplexAppendPlan

PersonaPlexPcmAppendBuffer

Transactionally frame 24 kHz float PCM into PersonaPlex 80 ms units.

pending_byte_count property

pending_byte_count: int

clear

clear() -> None

clear_force_listen

clear_force_listen() -> None

flush

flush(*, chunk_period_ms: int) -> dict[str, object] | None

has_pending

has_pending() -> bool

has_reserved

has_reserved() -> bool

prepare_append

prepare_append(
    payload: dict[str, object],
    *,
    operation_id: str,
    chunk_period_ms: int,
    allow_emit: bool,
) -> PersonaPlexPcmAppendReservation | None

prepare_commit

prepare_commit(
    *, operation_id: str, chunk_period_ms: int
) -> PersonaPlexPcmAppendReservation

PersonaPlexServingRuntimeAdapter

adapter_id class-attribute instance-attribute

adapter_id = 'personaplex'

clean_response_done_prefix class-attribute instance-attribute

clean_response_done_prefix = ''

data_plane instance-attribute

data_plane = PersonaPlexDataPlaneSession(encode_audio)

interrupted_tts_prefix class-attribute instance-attribute

interrupted_tts_prefix = ''

private_runtime_config_keys class-attribute instance-attribute

private_runtime_config_keys = _PRIVATE_RUNTIME_CONFIG_KEYS

session_states instance-attribute

session_states: dict[
    str, PersonaPlexServingSessionState
] = {}

capabilities staticmethod

capabilities(*, max_sessions: int) -> DuplexCapabilities

create_session_state

create_session_state() -> PersonaPlexServingSessionState

data_plane_context staticmethod

data_plane_context(
    *,
    epoch: int,
    turn_id: int,
    active_response_turn_id: int | None,
    active_response_id: str | None,
    auto_responds: bool,
    response_format: str,
    speed: float | None,
    modalities: tuple[str, ...],
) -> PersonaPlexDataPlaneContext

is_enabled staticmethod

is_enabled(config: object) -> bool

prepare_runtime_config async classmethod

prepare_runtime_config(
    config: object, *, model_config: Any
) -> dict[str, object]

remove_session_state

remove_session_state(session_id: str) -> None

runtime_config_for_update classmethod

runtime_config_for_update(
    config: object, current: Mapping[str, object]
) -> dict[str, object]

session_state

session_state(
    session_id: str,
) -> PersonaPlexServingSessionState

validate_client_extra_body staticmethod

validate_client_extra_body(extra_body: object) -> None

PersonaPlexStage0DuplexRuntime

Own per-session streaming Mimi encoders and first-append prefill.

device instance-attribute

device = device

max_sessions instance-attribute

max_sessions = max_sessions

model_path instance-attribute

model_path = model_path

request_sessions instance-attribute

request_sessions: dict[str, tuple[str, int]] = {}

sessions instance-attribute

stage_model instance-attribute

stage_model = stage_model

close_request

close_request(request_id: str) -> None

close_session

close_session(session_id: str, incarnation: int) -> None

prepare_append

prepare_append(
    duplex: dict[str, Any],
    *,
    prompt_len: int,
    request_id: str | None = None,
) -> PersonaPlexStage0PreparedAppend

record_sample

record_sample(
    *, request_id: str, text_token: Any, agent_codes: Any
) -> None

Commit one sampled temporal frame for the next live append.

PrefillStep dataclass

One tick of a recycled slot's system-prompt replay (see batched serving).

embedding class-attribute instance-attribute

embedding: Any = None

kind instance-attribute

kind: str

moshi_tokens class-attribute instance-attribute

moshi_tokens: Any = None

text_token class-attribute instance-attribute

text_token: int | None = None

user_sine class-attribute instance-attribute

user_sine: bool = False