vllm_omni.model_executor.models.gepard.nanocodec ¶
NeMo NanoCodec wrapper for Gepard — per-dimension FSQ decode.
Gepard's 32 FSQ heads emit per-dimension codes (8 groups x 4 dims), while stock NeMo decode wants composed per-group mixed-radix indices. decode_from_codes dequantizes per-dim instead and hands the result to stock decode_audio — a few lines on public NeMo API, no fork.
NeMo is a heavy optional dependency, imported lazily inside load. Decode only; encode_to_dequantized belongs to the voice-cloning follow-up.
END_OF_SPEECH_FADE_MS module-attribute ¶
END_OF_SPEECH_SILENCE_MS module-attribute ¶
LOOKBACK_FRAMES module-attribute ¶
NanoCodec ¶
Bases: Module
Thin wrapper owning the (optional) NeMo NanoCodec decoder.
Constructing this is cheap and NeMo-free; load does the heavy import + weight load so load_format=dummy and NeMo-less environments don't pay for it at model construction.
decode_from_codes ¶
(B, 32, T) per-dim codes -> waveform (B, T_audio). Returns audio only.
decode_stream ¶
decode_stream(
frames: Tensor,
*,
start_idx: int,
lookback: int = LOOKBACK_FRAMES,
is_final: bool = False,
) -> Tensor | None
Decode the tail of a (T, 32) code history past start_idx.
The window starts lookback frames earlier for receptive-field context and those samples are trimmed after, so successive calls concatenate to a single whole-history decode. None if nothing is new.
load ¶
Import NeMo, build the UnfoldedCodecModel, move to device, eval().
apply_end_of_speech_tail ¶
apply_end_of_speech_tail(
audio: Tensor,
sample_rate: int,
fade_ms: float = END_OF_SPEECH_FADE_MS,
silence_ms: float = END_OF_SPEECH_SILENCE_MS,
) -> Tensor
Fade the final fade_ms to zero, then append silence_ms of pad.
The fade is what removes the end-of-clip click; padding alone does not.