Skip to content

vllm_omni.model_executor.models.gepard.configuration_gepard

Config for Gepard-1.0, a single-stage autoregressive TTS.

Text tokens -> one 32-code FSQ audio frame per step -> NeMo NanoCodec -> waveform. The backbone is a vLLM-native Qwen3_5ForCausalLM; the codebook heads, binary stop head and voice-clone ref_compressor are Gepard additions.

Parses the model's gepard_config.json sidecar, which nests the LM parameters under backbone_config and carries audio-head cardinalities, special tokens, codec settings and the short-text repetition layout.

logger module-attribute

logger = logging.get_logger(__name__)

GepardConfig

Bases: PretrainedConfig

Configuration for the Gepard-1.0 native-AR TTS model.

Args mirror gepard_config.json. Defaults match the trained nineninesix/gepard-1.0 checkpoint so an instance built with no arguments (e.g. dummy/profiling loads) is still self-consistent.

audio_embed_dim instance-attribute

audio_embed_dim = audio_embed_dim

audio_head_levels instance-attribute

audio_head_levels = self._parse_audio_heads(audio_heads)

backbone_config instance-attribute

backbone_config = self._normalize_backbone(backbone_config)

codec_do_unfold instance-attribute

codec_do_unfold = cc.get('do_unfold', True)

codec_frame_rate_hz instance-attribute

codec_frame_rate_hz = cc.get('frame_rate_hz', 21.5)

codec_id instance-attribute

codec_id = cc.get(
    "codec_id",
    "nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps",
)

codec_sample_rate instance-attribute

codec_sample_rate = cc.get('sample_rate', 22050)

end_of_speech instance-attribute

end_of_speech = st.get('end_of_speech', 248071)

end_of_text instance-attribute

end_of_text = st.get('end_of_text', 248074)

fsq_levels instance-attribute

fsq_levels = cc.get('fsq_levels', [8, 7, 6, 6])

head0_vocab_size instance-attribute

head0_vocab_size = (
    self.audio_head_levels[0]
    if self.audio_head_levels
    else 8
)

keys_to_ignore_at_inference class-attribute instance-attribute

keys_to_ignore_at_inference = ['past_key_values']

model_type class-attribute instance-attribute

model_type = 'gepard'

num_audio_heads instance-attribute

num_audio_heads = len(self.audio_head_levels)

num_codec_groups instance-attribute

num_codec_groups = cc.get('num_layers', 8)

num_speaker_prefix instance-attribute

num_speaker_prefix = comp.get('num_queries', 8)

ref_compressor_d_model instance-attribute

ref_compressor_d_model = comp.get('d_model', 1024)

ref_compressor_ffn_mult instance-attribute

ref_compressor_ffn_mult = comp.get(
    "ffn_hidden_size_multiplier", 4
)

ref_compressor_num_blocks instance-attribute

ref_compressor_num_blocks = comp.get('num_layers', 2)

ref_compressor_num_heads instance-attribute

ref_compressor_num_heads = comp.get('num_heads', 8)

speaker_token_base instance-attribute

speaker_token_base = self.tokeniser_length

start_of_speech instance-attribute

start_of_speech = st.get('start_of_speech', 248070)

start_of_text instance-attribute

start_of_text = st.get('start_of_text', 248073)

stop_loss_weight instance-attribute

stop_loss_weight = stop_loss_weight

stop_pos_weight instance-attribute

stop_pos_weight = stop_pos_weight

stop_threshold instance-attribute

stop_threshold = stop_threshold

stop_token instance-attribute

stop_token = self.head0_vocab_size

temperature instance-attribute

temperature = temperature

text_repetition_apply_below instance-attribute

text_repetition_apply_below = tr.get('apply_below', 13)

text_repetition_enabled instance-attribute

text_repetition_enabled = tr.get('enabled', True)

text_repetition_max_repeats instance-attribute

text_repetition_max_repeats = tr.get('max_repeats', 8)

text_repetition_target_tokens instance-attribute

text_repetition_target_tokens = tr.get(
    "target_text_tokens", 16
)

tokeniser_length instance-attribute

tokeniser_length = st.get('tokeniser_length', 248077)

tts_pad instance-attribute

tts_pad = st.get('tts_pad', 248076)

voice_cloning_enabled instance-attribute

voice_cloning_enabled = vc.get('enabled', True)

from_checkpoint classmethod

from_checkpoint(
    model: str,
    backbone_config: dict | None = None,
    revision: str | None = None,
) -> GepardConfig

Build the full config for a checkpoint that self-identifies as the bare backbone: audio fields from the sidecar, backbone fields from the loaded config.

revision must be the one the weights came from — a revision that moves the audio-head cardinalities or the special tokens moves the prompt layout and the STOP sentinel with them.

get_text_config

get_text_config(**kwargs) -> PretrainedConfig

Return the Qwen3.5 backbone config, which vLLM runs the backbone on.