vllm_omni.model_executor.models.common.cfg_pairing ¶
Model-neutral classifier-free-guidance request pairing for the vLLM v1 scheduler.
A guided request decodes together with an unconditional companion so their logits can be blended every step. Blending is only correct when both members decode the same position in the same engine step, so :func:apply_cfg_patches makes the scheduler pair-aware: a lone member waits for its partner, partners stay adjacent in the waiting queue, their num_computed_tokens are equalized after every schedule, and they finish together.
Requests opt in through SamplingParams.extra_args::
cond: {"cfg_role": "cond", "cfg_pair_id": <id>}
uncond: {"cfg_role": "uncond", "cfg_pair_id": <id>}
cfg_pair_id is optional. A model whose serving layer cannot mint a pair id (the companion request id is only known inside the engine) may send cfg_role alone; the pair id is then derived from the request id, stripping the companion's registered suffix (see :func:register_cfg_uncond_suffix).
Logit blending itself is the caller's concern: this module only keeps the two rows step-locked. The patch is a no-op for engines that never see a CFG request (the pair registry stays empty).
apply_cfg_patches ¶
Make the vLLM v1 scheduler CFG-pair-aware. Idempotent, per process.
Must run in the engine-core process before Scheduler is constructed (the __init__ wrapper installs the pair registry).
cfg_patches_applied ¶
cfg_patches_applied() -> bool
True once :func:apply_cfg_patches has patched the scheduler in this process.
cfg_uncond_suffixes ¶
Companion request-id suffixes registered in this process.