Configuration Options¶
This section lists the most common options for running vLLM-Omni.
For options within a vLLM Engine, please refer to the vLLM 0.29 configuration guide.
Each model defines fixed topology in a registered PipelineConfig and runtime overrides in a deploy YAML.
For process-level settings shared by the CLI, deploy YAML, workers, and model integrations, see Environment Variables.
For a specific example, see the Qwen2.5-Omni deploy config. The matching frozen pipeline topology lives at vllm_omni/model_executor/models/qwen2_5_omni/pipeline.py.
For an introduction, see Pipeline and deploy configurations.
Memory Configuration¶
- GPU Memory Calculation and Configuration - Guide on how to calculate memory requirements and set up
gpu_memory_utilizationfor optimal performance
Optimization Features¶
- Diffusion Features Overview - Complete overview of all diffusion model features and supported models