vllm_omni.model_executor.models.audex.audex_omni ¶
Audex audio-understanding model (checkpoint_folder_full): S2T / audio QA.
Mirrors the Qwen2-Audio pattern with the 2B dense backbone
language_model= the vllm-omniNemotronDenseForCausalLMport (wrapped viainit_vllm_registered_model; resolved through the omni model registry).audio_tower= transformersQwen2AudioEncoder("NV-Whisper"); a parameter-free avg pool maps 1500 → 750 frames per padded 30 s clip.audio_projector= RMSNorm → fc1 → relu² → fc2 (no bias), matching the releaseNemotronHAudexProjector.
embed_multimodal(...) runs encoder+projector and returns one embedding tensor per audio item; vLLM's SupportsMultiModal.embed_input_ids merges them into the <so_embedding> placeholder positions. The encoder runs in microbatches over clips so peak memory is bounded regardless of audio length (up to the 900 s / 30-clip cap enforced by the processor).
AudexAudioFeatureInputs ¶
Bases: TensorSchema
Dimensions
- c: total number of 30s clips (flattened across audio items)
- b: number of audio items
AudexProjector ¶
NemotronDenseAudexForConditionalGeneration ¶
Bases: Module, SupportsMultiModal, SupportsPP
audio_projector instance-attribute ¶
audio_projector = AudexProjector(
audio_encoder_hidden_size=config.audio_encoder_hidden_size,
intermediate_size=config.audio_projector_intermediate_size,
text_hidden_size=config.hidden_size,
activation=config.audio_projector_activation,
norm_eps=config.audio_projector_norm_eps,
).to(llm_dtype)
audio_tower instance-attribute ¶
language_model instance-attribute ¶
language_model = init_vllm_registered_model(
vllm_config=vllm_config,
hf_config=config,
prefix=maybe_prefix(prefix, "language_model"),
architectures=[self._LM_ARCHITECTURE],
)
make_empty_intermediate_tensors instance-attribute ¶
forward ¶
forward(
input_ids: Tensor | None,
positions: Tensor,
intermediate_tensors: IntermediateTensors | None = None,
inputs_embeds: Tensor | None = None,
**kwargs: object,
) -> Tensor | IntermediateTensors
load_weights ¶
Split the combined weight stream into LLM / audio_tower / projector.
vLLM streams all .safetensors in the model dir; route audio_encoder.*/audio_projector.* to the audio modules and leave the remaining model.*/lm_head.* names untouched so the wrapped NemotronDenseForCausalLM loader applies its own qkv fusion and remapping.
NemotronHAudexForConditionalGeneration ¶
Bases: NemotronDenseAudexForConditionalGeneration, IsHybrid, HasInnerState, SupportsMambaPrefixCaching
Audex 30B-A3B audio-understanding model (checkpoint_folder_full).
Mirrors nvidia/Nemotron-Labs-Audex-30B-A3B's inference_scripts_vllm/audioqa_scripts/audex_30b_a3b_vllm/modeling_audex_vllm.py: the same encoder/projector/processing as the dense 2B wrapper around vLLM's native hybrid Mamba+MoE NemotronHForCausalLM, whose state hooks are delegated below.