vllm_omni.model_executor.models.gepard.configuration_gepard ¶
Config for Gepard-1.0, a single-stage autoregressive TTS.
Text tokens -> one 32-code FSQ audio frame per step -> NeMo NanoCodec -> waveform. The backbone is a vLLM-native Qwen3_5ForCausalLM; the codebook heads, binary stop head and voice-clone ref_compressor are Gepard additions.
Parses the model's gepard_config.json sidecar, which nests the LM parameters under backbone_config and carries audio-head cardinalities, special tokens, codec settings and the short-text repetition layout.
GepardConfig ¶
Bases: PretrainedConfig
Configuration for the Gepard-1.0 native-AR TTS model.
Args mirror gepard_config.json. Defaults match the trained nineninesix/gepard-1.0 checkpoint so an instance built with no arguments (e.g. dummy/profiling loads) is still self-consistent.
codec_id instance-attribute ¶
head0_vocab_size instance-attribute ¶
keys_to_ignore_at_inference class-attribute instance-attribute ¶
ref_compressor_ffn_mult instance-attribute ¶
ref_compressor_num_blocks instance-attribute ¶
text_repetition_apply_below instance-attribute ¶
text_repetition_max_repeats instance-attribute ¶
text_repetition_target_tokens instance-attribute ¶
from_checkpoint classmethod ¶
from_checkpoint(
model: str,
backbone_config: dict | None = None,
revision: str | None = None,
) -> GepardConfig
Build the full config for a checkpoint that self-identifies as the bare backbone: audio fields from the sidecar, backbone fields from the loaded config.
revision must be the one the weights came from — a revision that moves the audio-head cardinalities or the special tokens moves the prompt layout and the STOP sentinel with them.
get_text_config ¶
Return the Qwen3.5 backbone config, which vLLM runs the backbone on.