Skip to content

vllm_omni.model_extras.auk

Request helpers for AuK: prompt rendering, sampling params, duration policy.

AuK's contract is one instruction plus an optional reference clip in, audio out; the task (zero-shot TTS, instruct TTS, content or acoustic editing, enhancement, separation) is carried by the instruction text alone, so there is no per-task request surface.

DEFAULT_CFG module-attribute

DEFAULT_CFG = 2.0

DEFAULT_NFE module-attribute

DEFAULT_NFE = 32

DEFAULT_SWAY module-attribute

DEFAULT_SWAY = -1.0

HOP module-attribute

HOP = 480

NO_PROMPT_AUDIO module-attribute

NO_PROMPT_AUDIO = '|<no_prompt_audio>|'

SAMPLE_RATE module-attribute

SAMPLE_RATE = 24000

auk_prompt

auk_prompt(
    instruction: str,
    audio: tuple[ndarray | Tensor, int] | None = None,
    *,
    gen_seconds: float | None = None,
    sway: float = DEFAULT_SWAY,
    t_grid: list[float] | None = None,
    vae_sample: bool = False,
) -> dict[str, Any]

Build the stage-0 prompt for one AuK request.

Parameters:

Name Type Description Default
instruction str

The task instruction; for text-only requests the no-reference marker is appended if it is not already there.

required
audio tuple[ndarray | Tensor, int] | None

Reference or source clip as (waveform, sample_rate). None selects the text-only (instruct TTS) path.

None
gen_seconds float | None

Target duration. None reuses the source clip's length and is an error for text-only requests.

None
sway float

Sway coefficient for the Euler time grid.

DEFAULT_SWAY
t_grid list[float] | None

Explicit time grid, overriding sway.

None
vae_sample bool

Draw the reference latent from the VAE posterior instead of taking its mean.

False

Returns:

Type Description
dict[str, Any]

A prompt dict for Omni.generate.

auk_sampling_params

auk_sampling_params(
    *,
    nfe: int = DEFAULT_NFE,
    cfg: float = DEFAULT_CFG,
    seed: int | None = 0,
) -> list[SamplingParams | OmniDiffusionSamplingParams]

Build the per-stage sampling params for one AuK request.

Stage 0 is a prefill-only encoder, so its vLLM sampling is a no-op that only has to satisfy the scheduler. Stage 1 carries the real knobs: nfe Euler steps and CFG strength. AuK-Flash ignores both and pins its distilled recipe.

resolve_gen_frames

resolve_gen_frames(
    gen_seconds: float | None,
    ref_frames: int,
    *,
    sample_rate: int = SAMPLE_RATE,
    hop: int = HOP,
) -> int

Resolve the target latent length in frames.

An explicit duration wins; otherwise the target is as long as the source clip. A text-only request with no duration has nothing to generate, which is an error rather than a silent one-frame output.