Skip to content

vllm_omni.model_executor.models.audex.cfg

Classifier-free guidance for the Audex thinker stage.

CFG pairs a conditional request with an unconditional (null-prompt) companion in the same engine and blends their logits every step:

blended = uncond + cfg_scale * (cond - uncond)

Both rows receive the blended logits and the sampled token is copied from the cond row to the uncond row, so the two sequences stay token-identical. Requests opt in via SamplingParams.extra_args:

cond:   {"cfg_scale": 1.5, "cfg_role": "cond",   "cfg_pair_id": <id>}
uncond: {"cfg_scale": 1.5, "cfg_role": "uncond", "cfg_pair_id": <id>}

Blending is only correct when both pair members decode the same position in the same engine step. Keeping the pair step-locked is model-neutral and lives in :mod:vllm_omni.model_executor.models.common.cfg_pairing; the scheduler patches are re-exported here so existing imports keep working.

logger module-attribute

logger = init_logger(__name__)

AudexCFGLogitsProcessor

Bases: LogitsProcessor

Blend cond/uncond logits for classifier-free guidance.

Pairs are matched by cfg_pair_id. For each pair the blended logits are written to both rows so the sampler picks the same token; the post-sampling copy of the cond token into the uncond slot is installed by :meth:_ensure_sample_patched (which must run inside each worker process, hence from __init__ rather than a main-process patch).

apply

apply(logits: Tensor) -> Tensor

is_argmax_invariant

is_argmax_invariant() -> bool

update_state

update_state(batch_update: BatchUpdate | None) -> None

validate_params classmethod

validate_params(params: SamplingParams) -> None

apply_cfg_patches

apply_cfg_patches() -> None

Make the vLLM v1 scheduler CFG-pair-aware. Idempotent, per process.

Must run in the engine-core process before Scheduler is constructed (the __init__ wrapper installs the pair registry).

cfg_patches_applied

cfg_patches_applied() -> bool

True once :func:apply_cfg_patches has patched the scheduler in this process.