Skip to content

vllm_omni.model_executor.models.auk

Modules:

Name Description
auk

AuK (Tencent): instruction-driven speech generation and editing.

pipeline

AuK pipeline topology (frozen).

AuKForConditionalGeneration

Bases: Module, SupportsMultiModal, SupportsPP, SupportsMRoPE

Encoder stage of AuK: Qwen2.5-Omni thinker with learned layer fusion.

The forward pass re-implements the thinker language model's layer loop so the per-layer outputs can be fused on the fly: vLLM's decoder layers carry the residual stream separately, and hidden + residual after layer i is the HF output_hidden_states[i + 1] entry, while the last entry is the final-normed state. Only one accumulator lives at a time.

has_preprocess instance-attribute

has_preprocess = False

have_multimodal_outputs class-attribute instance-attribute

have_multimodal_outputs = True

make_empty_intermediate_tensors instance-attribute

make_empty_intermediate_tensors = (
    self.thinker.make_empty_intermediate_tensors
)

model instance-attribute

model = self.thinker

model_stage instance-attribute

model_stage = (
    getattr(vllm_config.model_config, "model_stage", None)
    or "encoder"
)

omni_pooler_payload_include_hidden class-attribute instance-attribute

omni_pooler_payload_include_hidden = False

requires_full_prefix_cached_hidden_states class-attribute instance-attribute

requires_full_prefix_cached_hidden_states = False

sampler cached property

sampler

thinker instance-attribute

thinker = init_vllm_registered_model(
    vllm_config=vllm_config,
    prefix=maybe_prefix(prefix, "thinker"),
    hf_config=thinker_config,
    architectures=["Qwen2_5OmniThinkerModel"],
)

thinker_config instance-attribute

thinker_config = thinker_config

vllm_config instance-attribute

vllm_config = vllm_config

compute_logits

compute_logits(
    hidden_states: Tensor | OmniOutput, **kwargs: object
) -> Tensor | None

Emit EOS immediately; the useful result is the prefill text condition.

embed_input_ids

embed_input_ids(
    input_ids: Tensor,
    multimodal_embeddings=None,
    is_multimodal=None,
) -> Tensor

embed_multimodal

embed_multimodal(**kwargs)

forward

forward(
    input_ids: Tensor,
    positions: Tensor,
    intermediate_tensors: IntermediateTensors | None = None,
    inputs_embeds: Tensor | None = None,
    **kwargs: object,
) -> OmniOutput

get_language_model

get_language_model() -> Module

get_mrope_input_positions

get_mrope_input_positions(
    input_tokens: list[int],
    mm_features: list[MultiModalFeatureSpec] | None = None,
    *,
    hf_config: PretrainedConfig,
    image_grid_thw: list[list[int]] | Tensor,
    video_grid_thw: list[list[int]] | Tensor,
    second_per_grid_ts: list[float] | None = None,
    context_len: int = 0,
    seq_len: int | None = None,
    audio_feature_lengths: Tensor | None = None,
    use_audio_in_video: bool = False,
) -> tuple[Tensor, int]

get_placeholder_str classmethod

get_placeholder_str(modality: str, i: int) -> str | None

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]