Skip to content

Autoregressive runtime

The AR runtime extends vLLM scheduling and worker execution for omni-stage inputs and outputs while preserving vLLM scheduling and cache semantics.

Candidate invariants

AR-INV-001: vLLM owns base scheduling semantics

Rule: Omni schedulers MUST preserve upstream request-state and cache transitions unless an Omni-specific difference is documented and tested.

AR-INV-002: Omni data crosses explicit adapters

Rule: Modality-specific stage data MUST be converted at an input or output adapter, not injected through unrelated scheduler state.

AR-INV-003: Workers execute assigned work

Rule: Workers and model runners MUST NOT implement cross-stage routing.

Safe-change guide

Test request lifecycle, abort, cache state, and every affected worker execution mode against the supported upstream vLLM contract.