vllm_omni.model_executor.models.audex.cfg ¶
Classifier-free guidance for the Audex thinker stage.
CFG pairs a conditional request with an unconditional (null-prompt) companion in the same engine and blends their logits every step:
blended = uncond + cfg_scale * (cond - uncond)
Both rows receive the blended logits and the sampled token is copied from the cond row to the uncond row, so the two sequences stay token-identical. Requests opt in via SamplingParams.extra_args:
cond: {"cfg_scale": 1.5, "cfg_role": "cond", "cfg_pair_id": <id>}
uncond: {"cfg_scale": 1.5, "cfg_role": "uncond", "cfg_pair_id": <id>}
Blending is only correct when both pair members decode the same position in the same engine step. Keeping the pair step-locked is model-neutral and lives in :mod:vllm_omni.model_executor.models.common.cfg_pairing; the scheduler patches are re-exported here so existing imports keep working.
AudexCFGLogitsProcessor ¶
Bases: LogitsProcessor
Blend cond/uncond logits for classifier-free guidance.
Pairs are matched by cfg_pair_id. For each pair the blended logits are written to both rows so the sampler picks the same token; the post-sampling copy of the cond token into the uncond slot is installed by :meth:_ensure_sample_patched (which must run inside each worker process, hence from __init__ rather than a main-process patch).
apply_cfg_patches ¶
Make the vLLM v1 scheduler CFG-pair-aware. Idempotent, per process.
Must run in the engine-core process before Scheduler is constructed (the __init__ wrapper installs the pair registry).