Skip to content

vllm_omni.model_executor.models.gepard.nanocodec

NeMo NanoCodec wrapper for Gepard — per-dimension FSQ decode.

Gepard's 32 FSQ heads emit per-dimension codes (8 groups x 4 dims), while stock NeMo decode wants composed per-group mixed-radix indices. decode_from_codes dequantizes per-dim instead and hands the result to stock decode_audio — a few lines on public NeMo API, no fork.

NeMo is a heavy optional dependency, imported lazily inside load. Decode only; encode_to_dequantized belongs to the voice-cloning follow-up.

END_OF_SPEECH_FADE_MS module-attribute

END_OF_SPEECH_FADE_MS = float(
    os.environ.get("VLLM_GEPARD_END_FADE_MS", "0.0")
)

END_OF_SPEECH_SILENCE_MS module-attribute

END_OF_SPEECH_SILENCE_MS = float(
    os.environ.get("VLLM_GEPARD_END_SILENCE_MS", "0.0")
)

LOOKBACK_FRAMES module-attribute

LOOKBACK_FRAMES = int(
    os.environ.get("VLLM_GEPARD_LOOKBACK_FRAMES", "8")
)

logger module-attribute

logger = init_logger(__name__)

NanoCodec

Bases: Module

Thin wrapper owning the (optional) NeMo NanoCodec decoder.

Constructing this is cheap and NeMo-free; load does the heavy import + weight load so load_format=dummy and NeMo-less environments don't pay for it at model construction.

codec_id instance-attribute

codec_id = codec_id

is_loaded property

is_loaded: bool

sample_rate instance-attribute

sample_rate = sample_rate

decode_from_codes

decode_from_codes(
    codes: Tensor, codes_len: Tensor
) -> Tensor

(B, 32, T) per-dim codes -> waveform (B, T_audio). Returns audio only.

decode_stream

decode_stream(
    frames: Tensor,
    *,
    start_idx: int,
    lookback: int = LOOKBACK_FRAMES,
    is_final: bool = False,
) -> Tensor | None

Decode the tail of a (T, 32) code history past start_idx.

The window starts lookback frames earlier for receptive-field context and those samples are trimmed after, so successive calls concatenate to a single whole-history decode. None if nothing is new.

load

load(device: device) -> None

Import NeMo, build the UnfoldedCodecModel, move to device, eval().

apply_end_of_speech_tail

apply_end_of_speech_tail(
    audio: Tensor,
    sample_rate: int,
    fade_ms: float = END_OF_SPEECH_FADE_MS,
    silence_ms: float = END_OF_SPEECH_SILENCE_MS,
) -> Tensor

Fade the final fade_ms to zero, then append silence_ms of pad.

The fade is what removes the end-of-clip click; padding alone does not.