vllm_omni.model_executor.models.gepard.prompt ¶
Gepard-1.0 prompt layout.
preprocess only consumes this layout — it injects the speaker prefix at the placeholder slots and samples the first audio frame from the SOS position. The layout itself is built here, by whichever entry point is assembling a request.
build_gepard_prompt_ids ¶
build_gepard_prompt_ids(
text_token_ids: list[int], *, config: GepardConfig
) -> list[int]
Assemble [speaker_slots] + [SOT, *text, EOT]*(R-1) + [SOT, *text, EOT, SOS].
Only the final copy carries SOS, the learned "audio starts now" gate, so the earlier copies are read as context and never voiced. Short texts repeat because a couple of text tokens carry too little mass against the speaker prefix. This must match the training layout or WER collapses, which is why the thresholds come from the checkpoint's text_repetition block rather than from literals here.
Raises ValueError on empty text: the layout would still be structurally valid and the model would voice it as arbitrary audio, so an empty request has to fail loudly. This is the only entry point in the offline path — there is no adapter above it to check first.