vllm_omni.model_extras.auk ¶
Request helpers for AuK: prompt rendering, sampling params, duration policy.
AuK's contract is one instruction plus an optional reference clip in, audio out; the task (zero-shot TTS, instruct TTS, content or acoustic editing, enhancement, separation) is carried by the instruction text alone, so there is no per-task request surface.
auk_prompt ¶
auk_prompt(
instruction: str,
audio: tuple[ndarray | Tensor, int] | None = None,
*,
gen_seconds: float | None = None,
sway: float = DEFAULT_SWAY,
t_grid: list[float] | None = None,
vae_sample: bool = False,
) -> dict[str, Any]
Build the stage-0 prompt for one AuK request.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
instruction | str | The task instruction; for text-only requests the no-reference marker is appended if it is not already there. | required |
audio | tuple[ndarray | Tensor, int] | None | Reference or source clip as | None |
gen_seconds | float | None | Target duration. | None |
sway | float | Sway coefficient for the Euler time grid. | DEFAULT_SWAY |
t_grid | list[float] | None | Explicit time grid, overriding | None |
vae_sample | bool | Draw the reference latent from the VAE posterior instead of taking its mean. | False |
Returns:
| Type | Description |
|---|---|
dict[str, Any] | A prompt dict for |
auk_sampling_params ¶
auk_sampling_params(
*,
nfe: int = DEFAULT_NFE,
cfg: float = DEFAULT_CFG,
seed: int | None = 0,
) -> list[SamplingParams | OmniDiffusionSamplingParams]
Build the per-stage sampling params for one AuK request.
Stage 0 is a prefill-only encoder, so its vLLM sampling is a no-op that only has to satisfy the scheduler. Stage 1 carries the real knobs: nfe Euler steps and CFG strength. AuK-Flash ignores both and pins its distilled recipe.
resolve_gen_frames ¶
resolve_gen_frames(
gen_seconds: float | None,
ref_frames: int,
*,
sample_rate: int = SAMPLE_RATE,
hop: int = HOP,
) -> int
Resolve the target latent length in frames.
An explicit duration wins; otherwise the target is as long as the source clip. A text-only request with no duration has nothing to generate, which is an error rather than a silent one-frame output.