vllm_omni.model_executor.models.audex.audex_code2wav ¶
Audex code2wav: streaming speech-codec-token → waveform stage.
Wraps the vendored Audex causal speech decoder (see speech_decoder/). Unlike the flow/vocoder stages of other TTS models, the decoder is natively stateful and streaming: one AudexCausalSpeechDecoderSession per request holds the decode cache, each engine step pushes only the newly received codec frames, and the session's 4-frame lookahead tail is drained by flush() when the request finishes. The stage therefore never re-decodes left context and always returns delta audio.
AudexCode2Wav ¶
Bases: Module
Stage-1 model for Audex TTS (GenerationModelRunner).
enable_update_additional_information instance-attribute ¶
compute_logits ¶
compute_logits(
hidden_states: Tensor | OmniOutput,
sampling_metadata: Any = None,
) -> None
forward ¶
forward(
input_ids: Tensor | None = None,
positions: Tensor | None = None,
intermediate_tensors: Any = None,
inputs_embeds: Tensor | None = None,
runtime_additional_information: list[dict[str, Any]]
| None = None,
**kwargs: Any,
) -> OmniOutput
make_omni_output ¶
make_omni_output(
model_outputs: Tensor | OmniOutput | tuple,
**kwargs: Any,
) -> OmniOutput