Skip to content

vllm_omni.model_executor.models.personaplex.personaplex_mimi

Streaming Mimi codec for PersonaPlex (moshi-free).

Frame-clocked duplex needs a codec that encodes/decodes exactly one 80 ms frame per call with state carried across the conversation. The moshi package provided that; this module removes the dependency by combining:

  • transformers MimiModel (kyutai/mimi — the same checkpoint family PersonaPlex ships) for the quantizer and every SEANet conv weight. Its streaming support only covers the encoder convs, so conv streaming is done here instead, uniformly.
  • Our own streaming wrappers mirroring the Moshi reference semantics (MIT), verified against recorded reference outputs:
  • Conv1d: left-context carry of effective_kernel - stride samples, zero-initialized at stream start (pad_mode="constant").
  • ConvTranspose1d: overlap-add tail carry of kernel - stride output samples, with the double-counted bias subtracted on merge.
  • transformer: the encoder/decoder transformers use a 250-position sliding context over a ring KV with absolute-offset RoPE. Hugging Face's cache path diverges once the window engages (position 250), so the transformers are reimplemented here on the same ring-KV design as personaplex_temporal.py and loaded directly from the PersonaPlex checkpoint's fused layout (LayerNorm + per-layer LayerScale + GELU FFN).

All per-stream state is [B, ...] with per-row reset (reset_slot), so the codec composes with elastic slot recycling in batched duplex serving.

CODEBOOKS module-attribute

CODEBOOKS = 8

DEFAULT_HF_REPO module-attribute

DEFAULT_HF_REPO = 'kyutai/mimi'

FRAME_SIZE module-attribute

FRAME_SIZE = 1920

PersonaPlexMimiCodec

Bases: Module

Streaming Mimi encode/decode at one 80 ms frame per call (moshi-free).

decoder_transformer instance-attribute

decoder_transformer = _MimiStreamingTransformer().to(
    self.device, self.dtype
)

device instance-attribute

device = torch.device(device)

dtype instance-attribute

dtype = next(self.model.parameters()).dtype

encoder_transformer instance-attribute

encoder_transformer = _MimiStreamingTransformer().to(
    self.device, self.dtype
)

model instance-attribute

model = self.model.to(self.device).eval()

decode_frame

decode_frame(
    codes: Tensor, active: Tensor | None = None
) -> Tensor

[B, 8] codes -> [B, frame_size] float PCM.

decode_frames

decode_frames(
    codes: Tensor, active: Tensor | None = None
) -> Tensor

[B, 8, F] codes -> [B, F * frame_size] float PCM.

encode_frame

encode_frame(
    pcm: Tensor, active: Tensor | None = None
) -> Tensor

[B, frame_size] float PCM -> [B, 8] codes.

reset_slot

reset_slot(b: int) -> None

reset_streaming

reset_streaming() -> None

streaming_init

streaming_init(batch_size: int) -> None