vllm_omni.model_executor.models.audex.audex_xcodec ¶
Audex TTA stage 1: XCodec1 RVQ codes → waveform.
Unlike the streaming causal speech decoder used for TTS, XCodec1 is a CNN codec that the official flow decodes over the full sequence, so this stage runs sync full-payload only: each request delivers its complete de-interleaved [frames, 4] codec-id payload once at finish and the decode happens in a single call.
The checkpoint (hf-audio/xcodec-hubert-general-balanced) is external to the Audex repo and loads through transformers remote code; the model skeleton is built from config at startup for vLLM's memory profiler and the weights arrive via the standard load_weights iteration over the snapshot's safetensors.
AudexXCodec1 ¶
Bases: Module
Stage-1 model for Audex TTA (GenerationModelRunner).
enable_update_additional_information instance-attribute ¶
compute_logits ¶
compute_logits(
hidden_states: Tensor | OmniOutput,
sampling_metadata: Any = None,
) -> None
forward ¶
forward(
input_ids: Tensor | None = None,
positions: Tensor | None = None,
intermediate_tensors: Any = None,
inputs_embeds: Tensor | None = None,
runtime_additional_information: list[dict[str, Any]]
| None = None,
**kwargs: Any,
) -> OmniOutput
make_omni_output ¶
make_omni_output(
model_outputs: Tensor | OmniOutput | tuple,
**kwargs: Any,
) -> OmniOutput