Design Documents¶
This section contains design documents and architecture specifications for vLLM-Omni. The sidebar groups documents by the system concern they describe; retired module pages are preserved in the legacy archive instead of remaining in the active navigation.
Architecture Documents¶
Feature Design Documents¶
For user-facing configuration and current compatibility, see the Features overview. A design document defines an implementation contract; it is not, by itself, a general support claim.
Runtime and stage execution¶
- Full-Duplex Runtime (MiniCPM-o 4.5)
- Full-Duplex Runtime (PersonaPlex)
- Disaggregated Inference
- Host Weight Runtime
- Async Chunk
- Async Diffusion Output
- Async Omni Output Materialization
- Runner-to-model Prefill/Decode Phase Contract
- Automatic Prefix Caching in Omni Models
- Model-local KV Caches
- Realtime AR-Diffusion Sessions
Communication¶
OmniConnector implementations¶
- Mooncake Store Connector
- Mooncake Transfer Engine Connector
- Mori Transfer Engine Connector
- NIXL Connector
- Shared Memory Connector
- Yuanrong Store Connector
- Yuanrong Transfer Engine Connector
Quantization¶
Diffusion acceleration¶
Parallelism¶
- CFG-Parallel
- Expert Parallel
- Hybrid Sharded Data Parallel (HSDP)
- Pipeline Parallel
- Sequence Parallel
- Tensor Parallel
- VAE Patch Parallelism
KV cache and memory management¶
Attention optimization¶
The Diffusion Attention Backends guides list selectable backends, platform defaults, installation, and tuning. The design contracts separate selection mechanics from backend algorithms:
CPU offloading¶
- Overview and Shared Contracts
- Model-Level Offload
- Layerwise Offload
- TeaCache
- Diffusion Continuous Batching
Infrastructure and Performance¶
Module Design Documents¶
- Entrypoints and Serving Boundaries
- vLLM-Omni Configuration
- Input, Output, and Modality Contracts
- Error Classification, Propagation, and Rendering
- Engine Orchestration
- Stage Runtime and Replica Lifecycle
- OmniConnector
- Model Integration
- Autoregressive Runtime
- Diffusion
- Execution Platforms
- Cache Management
- Host Weight Runtime
- Quantization
- Observability
- Profiling
- Benchmarking
The pre-#5137 pages are preserved in the legacy module archive for historical reference and are not active design contracts.