vllm_omni.model_executor.models.gepard.gepard_talker ¶
Gepard-1.0 native-AR talker.
Single-stage AR TTS on a vLLM-native Qwen3.5 backbone. Each step samples one 32-channel FSQ frame; head0 is the vLLM-facing token and the other 31 are side-channel. Generation ends on a learned binary stop head, not an EOS token. The NeMo NanoCodec decodes committed frames to a waveform outside vLLM.
Zero-shot only, and enforce_eager; CUDA graph is a perf follow-up.
CHUNK_FRAMES module-attribute ¶
FIRST_CHUNK_FRAMES module-attribute ¶
GepardTalkerForConditionalGeneration ¶
Bases: Module
Gepard native-AR TTS talker (Qwen3.5 backbone + 32 FSQ heads + stop).
audio_embed_proj instance-attribute ¶
audio_embed_proj = nn.Sequential(
nn.Linear(
self.num_heads * cfg.audio_embed_dim, hidden
),
nn.GELU(),
nn.Linear(hidden, hidden),
nn.LayerNorm(hidden, elementwise_affine=False),
)
audio_embeddings instance-attribute ¶
audio_embeddings = nn.ModuleList(
[
nn.Embedding(v, cfg.audio_embed_dim)
for v in self.vocab_sizes
]
)
fused_codebook_head instance-attribute ¶
model instance-attribute ¶
null_prefix instance-attribute ¶
compute_logits ¶
compute_logits(
hidden_states: Tensor | OmniOutput,
sampling_metadata: Any = None,
)
forward ¶
forward(
input_ids: Tensor,
positions: Tensor,
intermediate_tensors: IntermediateTensors | None = None,
inputs_embeds: Tensor | None = None,
**kwargs: Any,
) -> Tensor | OmniOutput
load_weights ¶
Fuse codebook_heads.{0..31} into fused_codebook_head; delegate the rest.
make_omni_output ¶
make_omni_output(
model_outputs: Tensor | OmniOutput, **kwargs: Any
) -> OmniOutput