vllm_omni.model_executor.models.personaplex.personaplex_mimi ¶
Streaming Mimi codec for PersonaPlex (moshi-free).
Frame-clocked duplex needs a codec that encodes/decodes exactly one 80 ms frame per call with state carried across the conversation. The moshi package provided that; this module removes the dependency by combining:
- transformers
MimiModel(kyutai/mimi— the same checkpoint family PersonaPlex ships) for the quantizer and every SEANet conv weight. Its streaming support only covers the encoder convs, so conv streaming is done here instead, uniformly. - Our own streaming wrappers mirroring the Moshi reference semantics (MIT), verified against recorded reference outputs:
Conv1d: left-context carry ofeffective_kernel - stridesamples, zero-initialized at stream start (pad_mode="constant").ConvTranspose1d: overlap-add tail carry ofkernel - strideoutput samples, with the double-counted bias subtracted on merge.- transformer: the encoder/decoder transformers use a 250-position sliding context over a ring KV with absolute-offset RoPE. Hugging Face's cache path diverges once the window engages (position 250), so the transformers are reimplemented here on the same ring-KV design as
personaplex_temporal.pyand loaded directly from the PersonaPlex checkpoint's fused layout (LayerNorm + per-layer LayerScale + GELU FFN).
All per-stream state is [B, ...] with per-row reset (reset_slot), so the codec composes with elastic slot recycling in batched duplex serving.
PersonaPlexMimiCodec ¶
Bases: Module
Streaming Mimi encode/decode at one 80 ms frame per call (moshi-free).