Skip to content

vllm_omni.model_executor.models.audex.speech_decoder

Vendored Audex causal speech decoder (Apache-2.0).

Verbatim copy of the audex_causal_speech_decoder remote-code package from nvidia/Nemotron-Labs-Audex-2B. Vendored because transformers' dynamic module loader resolves HuggingFace-cache blob symlinks to their real paths, which breaks the package's relative imports when loading with trust_remote_code directly from a hub snapshot.

Modules:

Name Description
configuration_audex_causal_speech_decoder
modeling_audex_causal_speech_decoder
streaming_utils

AudexCausalSpeechDecoderConfig

Bases: PretrainedConfig

codebook_levels instance-attribute

codebook_levels = codebook_levels or [
    4,
    4,
    4,
    4,
    4,
    4,
    4,
    4,
]

codebook_size instance-attribute

codebook_size = codebook_size

depth instance-attribute

depth = depth

embed_tokens_from_codes instance-attribute

embed_tokens_from_codes = embed_tokens_from_codes

heads instance-attribute

heads = heads

hidden_dim instance-attribute

hidden_dim = hidden_dim

hop_length instance-attribute

hop_length = hop_length

lookahead_steps instance-attribute

lookahead_steps = lookahead_steps

model_type class-attribute instance-attribute

model_type = 'audex_causal_speech_decoder'

pos_meb_dim instance-attribute

pos_meb_dim = pos_meb_dim

sample_rate instance-attribute

sample_rate = sample_rate

token_embed_dim instance-attribute

token_embed_dim = token_embed_dim

vq_dim instance-attribute

vq_dim = vq_dim

AudexCausalSpeechDecoderModel

Bases: PreTrainedModel

Cache class-attribute instance-attribute

all_tied_weights_keys class-attribute instance-attribute

all_tied_weights_keys: dict[str, Any] = {}

audex_speech_token_embedder instance-attribute

audex_speech_token_embedder = AudexSpeechTokenEmbedder(
    output_dim=config.vq_dim,
    token_embed_dim=config.token_embed_dim,
    codebook_levels=config.codebook_levels,
)

base_model_prefix class-attribute instance-attribute

base_model_prefix = 'module'

config_class class-attribute instance-attribute

lookahead_steps property

lookahead_steps: int

module instance-attribute

module = CausalCodecDecoderVocos(
    hidden_dim=config.hidden_dim,
    depth=config.depth,
    heads=config.heads,
    pos_meb_dim=config.pos_meb_dim,
    hop_length=config.hop_length,
    vq_dim=config.vq_dim,
    lookahead_steps=config.lookahead_steps,
)

create_cache

create_cache() -> CausalCodecDecoderCache

create_session

create_session(
    *,
    chunk_frames: int = 1,
    sample_rate: int | None = None,
    return_numpy: bool = True,
) -> AudexCausalSpeechDecoderSession

decode_cached

decode_cached(
    vq_emb: Tensor,
    cache: CausalCodecDecoderCache,
    lookahead_vq_emb: Tensor | None = None,
) -> Tensor

forward

forward(
    vq_emb: Tensor,
    patched_wav: Tensor | None = None,
    alpha: float = 0.0,
) -> Tensor

AudexCausalSpeechDecoderSession

buffer instance-attribute

buffer: list[list[int]] = []

cache instance-attribute

cache = decoder.create_cache()

chunk_frames instance-attribute

chunk_frames = chunk_frames

decoder instance-attribute

decoder = decoder

device property

device: device

return_numpy instance-attribute

return_numpy = return_numpy

sample_rate instance-attribute

sample_rate = sample_rate

flush

flush() -> Iterator[tuple[int, Any]]

push

push(
    token_frames: Sequence[Sequence[int]],
) -> Iterator[tuple[int, Any]]

reset

reset() -> None