vllm-omni serve¶
Multiple API frontends¶
For local EngineCore pipelines with one or more stages, --api-server-count starts multiple API frontend processes that share one parent-owned set of stage engines:
This mode is intended for workloads whose request parsing, multimodal input processing, or response encoding can bottleneck a single frontend. It currently requires local multi-process EngineCore stages with one local process group per replica. A stage may still use multiple num_replicas. It cannot be combined with diffusion, headless or remote stages, intra-stage data parallelism, Ray, fault tolerance, elastic expert parallelism, sleep mode, or runtime LoRA updating. Runtime voice upload and deletion are also disabled because those mutations are process-local; built-in voices, inline reference audio, and voices restored at startup remain available. The /v1/omni/sleep and /v1/omni/wakeup control routes return HTTP 409 for the same reason: their bookkeeping is process-local while stage engines are shared.
A single-stage model pipeline is supported through the normal launch command; this is distinct from the distributed --stage-id / --headless launch mode. Every stage in the pipeline must be an EngineCore stage. Mixed AR-to-DiT pipelines, including MiniMax-H3, remain unsupported even when their first stage is an EngineCore stage. Diffusion support requires per-frontend channels and shared-engine lifecycle handling in the diffusion backend.
Workloads and performance validation¶
Each API process owns its frontend processor, Orchestrator, request state, and per-replica EngineCore channels. Requests return through the channel of the originating API process; model weights and GPU stage engines are shared. This distributes frontend and orchestration work across CPU processes.
| Workload | Potential benefit | Limitation |
|---|---|---|
| Concurrent cold image/video inputs on supported EngineCore pipelines | Parallel frontend processing when one API process is CPU-limited | A single cold request and GPU computation may be unchanged |
| Many replicas producing frequent text/audio outputs | Less API/Orchestrator work and GIL contention per process | Routing remains serial within each Orchestrator |
| Warm inputs or low-concurrency, GPU-limited generation | Additional frontend capacity | More API processes may add overhead without improving throughput |
Here, a cold input means media or preprocessing caches miss, not that model weights must be loaded or GPU kernels compiled. Each frontend has its own processor/cache state, so cache duplication, CPU thread contention, and memory usage also matter. Multi-API serving does not itself change the output queue, move response encoding to an executor, or change the model's numerical scheduler.
Compare --api-server-count 1, 2, and higher counts with the same hardware, model/stage configuration, media, sampling parameters, output lengths, and offered load. Keep cold-cache and warmed-cache runs separate and report repeated measurements. Record CPU and request counts for every API worker, input and output queue delay, GPU utilization, first-output latency, and end-to-end latency. Confirm requests reach every worker: connection reuse can concentrate traffic on one process. Queue-only or single-API experiments do not establish a multi-API throughput gain. The multi-replica workload in #4680 remains a relevant validation target.
Stage-based CLI quickstart¶
The stage-based CLI is designed for deployments that require launching each pipeline stage in an isolated process (e.g., across separate operating system processes, distinct GPUs, or distributed hosts).
- For migrated models that utilize the bundled deployment YAML configurations located in
vllm_omni/deploy/, the--deploy-configflag is only required to override the default configuration. By default, executingvllm serve MODEL --omni ...automatically loads the bundled deployment configuration.
Example: Initializing Stage 0 (Orchestrator and API Server): The commands below show a common device mapping where Stage 0 uses GPU 0 and worker stages use GPU 1 via CUDA_VISIBLE_DEVICES.
CUDA_VISIBLE_DEVICES=0 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
--port 8091 \
--stage-id 0 \
--omni-master-address 127.0.0.1 \
--omni-master-port 26000
Example: Initializing a Headless Worker Stage (Stage 1):
CUDA_VISIBLE_DEVICES=1 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
--stage-id 1 \
--headless \
--omni-master-address 127.0.0.1 \
--omni-master-port 26000
When utilizing a custom deployment YAML, append --deploy-config /path/to/override.yaml to each command execution.
In the standard execution paradigm, the --stage-overrides argument is utilized to apply stage-specific configurations from a single CLI command. However, under the stage-based CLI paradigm, where each process strictly encapsulates a single stage, it is recommended to specify tuning parameters directly via discrete command-line flags for the respective stage, rather than constructing a composite --stage-overrides JSON string.
For example, as an alternative to the following composite configuration:
vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni --port 8091 \
--stage-overrides '{"1": {"gpu_memory_utilization": 0.5}}'
the stage-based CLI permits the direct initialization of Stage 1 with explicit parameters:
CUDA_VISIBLE_DEVICES=1 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
--stage-id 1 \
--headless \
--gpu-memory-utilization 0.5 \
--omni-master-address 127.0.0.1 \
--omni-master-port 26000
JSON CLI Arguments¶
Arguments¶
OmniConfig¶
Configuration for vLLM-Omni multi-stage and diffusion models.
--omni¶
Enable vLLM-Omni mode for multi-modal and diffusion models
Default: False
--enable-sleep-mode¶
Enable GPU memory pool for sleep mode.
Default: False
--task-type¶
Model-defined startup task type. The selected model validates supported values; for example, TTS models accept CustomVoice, VoiceDesign, or Base, while diffusion models may use it to select task-specific weights. If omitted, the model default is used.
Default: None
--forced-aligner¶
Enable streaming TTS word timestamps via a forced aligner. Pass the aligner model path/name, e.g. 'Qwen/Qwen3-ForcedAligner-0.6B'. Disabled when omitted.
Default: None
--forced-aligner-config¶
Optional YAML file for forced aligner settings (model, runner, gpu_memory_utilization, dtype, max_model_len). The --forced-aligner flag, when set, overrides the YAML model field.
Default: None
--forced-aligner-device¶
Device(s) for the forced-aligner stage (e.g. '2'). Defaults to sharing an existing stage's GPU when unset.
Default: None
--deploy-config¶
Path to a deploy config YAML (new format with stages/engine_args).
Default: None
--strategy-config¶
Path to a composable-parallel strategy.yaml. Only applies to registry-based models (e.g. qwen2_5_omni): its derived parallel sizing is overlaid onto the registry-merged stages before per-stage engine args are built.
Default: None
--stage-overrides¶
Per-stage JSON overrides. Example: '{"0": {"gpu_memory_utilization": 0.8}, "2": {"enforce_eager": true}}'
Default: None
--async-chunk, --no-async-chunk¶
Override the deploy YAML's async_chunk: bool. Unset leaves the YAML value in force.
Default: None
--stage-id¶
Select and launch a single stage by stage_id.
Default: None
--replica-id¶
Deprecated and ignored — replica ids are auto-assigned by the master server. Specifying this flag prints a warning and has no effect.
Default: None
--stage-init-timeout¶
The timeout for initializing a single stage in seconds (default: 300)
Default: 300
--init-timeout¶
The timeout for initializing the stages.
Default: 600
--shm-threshold-bytes¶
The threshold for the shared memory size.
Default: 65536
--log-stats¶
Enable logging the stats.
Default: False
--log-file¶
The path to the log file.
Default: None
--batch-timeout¶
The timeout for the batch.
Default: 10
--worker-backend¶
Possible choices: multi_process, ray
The backend to use for stage workers.
Default: multi_process
--ray-address¶
The address of the Ray cluster to connect to.
Default: None
--omni-master-address, -oma¶
Hostname or IP address of the Omni orchestrator (master).
Default: None
--omni-master-port, -omp¶
Port of the Omni orchestrator (master).
Default: None
--omni-replica-address, -ora¶
Local bind address (this host's IP) that the headless stage advertises to the Omni master for its handshake/input/output ZMQ sockets. If unset, auto-detected via a UDP-connect routing probe against --omni-master-address. Override only when the auto-detected IP is wrong (e.g. multi-NIC host where the master is reachable on the wrong interface).
Default: None
--omni-dp-size-local¶
Number of stage replicas this runtime launches locally for its own --stage-id. Process-local: head and every headless invocation read their own copy; values may differ across invocations. Requires --stage-id to be set when not equal to 1.
Default: 1
--omni-lb-policy¶
Possible choices: random, round-robin, least-queue-length
Per-stage load-balancing policy used by the head's StagePool to route requests across UP replicas. Only consulted on the head runtime.
Default: random
--omni-heartbeat-timeout¶
Seconds before an unreporting replica is marked ERROR in the OmniCoordinator. Only consulted on the head runtime.
Default: 30.0
--num-gpus¶
Number of GPUs to use for diffusion model inference.
Default: None
--model-class-name¶
Override the diffusion pipeline class name (e.g. LTX2Pipeline).
Default: None
--diffusion-load-format¶
Possible choices: default, custom_pipeline, dummy, diffusers
How to load the diffusion pipeline: native/registry (default), custom_pipeline, dummy, or diffusers for the HF diffusers adapter.
Default: None
--lora-path¶
LoRA checkpoint path(s) loaded when the diffusion server starts. The distill backend accepts one file per pipeline transformer.
Default: None
--lora-backend¶
Possible choices: peft, distill
Diffusion LoRA loading backend. 'distill' fuses checkpoint files into the base model at server startup; 'peft' uses the adapter manager.
Default: None
--lora-scale¶
Scale for a startup PEFT LoRA. Distilled LoRAs are fused at their checkpoint scale.
Default: None
--diffusion-compile-granularity¶
Possible choices: regional, full
Compilation scope for the generic diffusion model runner. 'regional' compiles repeated blocks (default); 'full' compiles the whole transformer and is incompatible with HSDP, sequence parallelism, CPU offload, and layerwise offload.
Default: None
--diffusion-compile-dynamic, --no-diffusion-compile-dynamic¶
Use dynamic shapes for the selected generic diffusion compile scope. Disable for fixed-shape workloads with --no-diffusion-compile-dynamic.
Default: None
--fa-deterministic¶
Request FlashAttention deterministic=True on the local FLASH_ATTN dense path. Slower than the library default deterministic=False; intended for accuracy CI. Serving default remains non-deterministic.
Default: False
--diffusers-load-kwargs¶
JSON object passed to DiffusionPipeline.from_pretrained().It overrides corresponding parameters in the standard vLLM-Omni interface.(e.g. '{"use_safetensors": true, "variant": "fp16"}').
Default: {}
--diffusers-call-kwargs¶
JSON object passed to pipeline.call(). Useful for model-specific sampling parameters not covered by the vLLM-Omni interface.During request time, it is overridden by corresponding parameters in the vLLM-Omni interface.(e.g. '{"num_inference_steps": 30, "guidance_scale": 7.5}').
Default: {}
--custom-pipeline-args¶
JSON object passed to native/custom diffusion pipelines. Only args containing "pipeline_class" trigger custom pipeline re-initialization.
Default: None
--usp, --ulysses-degree¶
Ulysses Sequence Parallelism degree for diffusion models. Equivalent to setting DiffusionParallelConfig.ulysses_degree.
Default: None
--ulysses-mode¶
Possible choices: strict, advanced_uaa
Ulysses sequence-parallel mode for diffusion models. 'strict' keeps the original divisibility requirements; 'advanced_uaa' enables the experimental UAA path for uneven sequence/head shapes.
Default: strict
--ulysses-a2a-permute, --no-ulysses-a2a-permute¶
Enable fused permute-free Ulysses all-to-all over NCCL symmetric memory. Only strict Ulysses layouts are eligible. Defaults to disabled.
Default: None
--ring, --ring-degree¶
Ring Sequence Parallelism degree for diffusion models. Equivalent to setting DiffusionParallelConfig.ring_degree.
Default: None
--allgather-degree¶
AllGather-KV Sequence Parallelism degree for non-causal diffusion attention. Equivalent to setting DiffusionParallelConfig.allgather_degree.
Default: None
--diffusion-quantization-config¶
JSON string for diffusion quantization_config. Example: '{"method":"fp8","activation_scheme":"dynamic"}'.
Default: None
--force-cutlass-fp8¶
Diffusion-only runtime override for ModelOpt FP8 checkpoints: force CUTLASS FP8 linear kernels on CUDA SM89+ devices. Ignored for BF16, non-ModelOpt FP8, ROCm, and older CUDA GPUs.
Default: None
--use-hsdp¶
Enable HSDP (Hybrid Sharded Data Parallel) for diffusion models. Shards model weights across GPUs to reduce per-GPU memory usage.
Default: False
--hsdp-shard-size¶
Number of GPUs to shard weights across. -1 = auto (world_size / replicate_size).
Default: -1
--hsdp-replicate-size¶
Number of replica groups for HSDP. Each group holds a full sharded copy.
Default: 1
--diffusion-attention-backend¶
Diffusion attention backend (shorthand). Sets the default backend for all diffusion attention roles, e.g. 'FLASH_ATTN'. May be combined with --diffusion-attention-config.per_role.* overrides, but mutually exclusive with --diffusion-attention-config.default.backend.
Default: None
--fastvideo-vsa-topk¶
Number of key/value blocks selected per query block by FASTVIDEO_VSA.
Default: None
--diffusion-attention-config, -dac¶
Diffusion attention config. Accepts JSON or vLLM-style dotted flags. Examples: --diffusion-attention-config.default.backend FLASH_ATTN, --diffusion-attention-config.default.backend TRTLLM_ATTN --diffusion-attention-config.default.skip_softmax.target_sparsity 0.5, --diffusion-attention-config.per_role.cross.backend SAGE_ATTN, --diffusion-attention-config '{"default": {"backend": "FLASH_ATTN"}, "per_role": {"cross": {"backend": "SAGE_ATTN"}}}'.
Default: None
--cache-backend¶
Cache backend for diffusion models, options: 'tea_cache', 'cache_dit', 'mag_cache', 'sea_cache', 'step_cache'
Default: none
--cache-config¶
JSON string of cache configuration. TeaCache: '{"rel_l1_thresh": 0.2}'. MagCache: '{"mag_threshold": 0.24, "mag_max_skip_steps": 5, "mag_retention_ratio": 0.1}'. Calibration mode: add '"mag_calibrate": true'
Default: None
--video-output-transport¶
JSON object configuring video output preparation, for example '{"enable_device_postprocess": true}'.
Default: None
--enable-cache-dit-summary¶
Enable cache-dit summary logging after diffusion forward passes.
Default: False
--step-execution¶
Enable per-step diffusion execution so running requests can be aborted between denoise steps.
Default: False
--request-batch-max-wait-ms¶
Request-mode batch admission: max milliseconds to wait for compatible requests to accumulate before scheduling a fused forward wave. 0 disables admission (default).
Default: 0.0
--vae-use-slicing¶
Enable VAE slicing for memory optimization (useful for mitigating OOM issues).
Default: False
--vae-use-tiling¶
Enable VAE tiling for memory optimization (useful for mitigating OOM issues).
Default: False
--vae-fast-path¶
Possible choices: off, lossless, channels_last
Wan VAE decoder fast path. 'lossless' (default) installs bit-exact fused kernels; 'channels_last' additionally switches decoder convolutions to channels-last memory format and fuses RMSNorm+SiLU (faster, not bit-exact); 'off' keeps the reference diffusers implementation.
Default: lossless
--disable-multithread-weight-load¶
Disable multi-threaded safetensors loading (default: enabled with 4 threads).
Default: True
--enable-broadcast-weight-load¶
Enable Rank-0 shared weight broadcast across workers for HSDP (default: disabled).
Default: False
--num-weight-load-threads¶
Number of threads for parallel weight loading (default: 4).
Default: 4
--diffusion-offload-config¶
Diffusion CPU-offload config as JSON. Set mode to module or layer, list dit and/or text_encoder in components, and put layer-only tuning under layer_options. Layer settings are weight_transfer (rank-local or allgather) and resident_layers (DiT only).
Default: None
--enable-cpu-offload¶
Compatibility alias for model-level CPU offload. New integrations should use --diffusion-offload-config with mode=module and explicit components.
Default: False
--enable-layerwise-offload¶
Compatibility alias for layerwise CPU offload. New integrations should use --diffusion-offload-config with mode=layer and explicit components.
Default: False
--enable-distributed-layerwise-offload¶
Compatibility alias for distributed layerwise CPU offload. New integrations should use mode=layer and configure weight transfer per component.
Default: False
--dlo-use-allgather¶
Compatibility option; use component weight_transfer=allgather in new configurations. Use shard + AllGather for weight reconstruction (default: True). When disabled (--dlo-no-use-allgather), each rank streams the standard loader's rank-local tensors via H2D only — no additional DP sharding, no AllGather, and no concurrent-request requirement.
Default: True
--dlo-no-use-allgather¶
Compatibility option; use component weight_transfer=rank-local in new configurations. Disable AllGather and stream standard-loader rank-local weights independently (including existing TP shards).
Default: True
--dlo-resident-layers¶
Compatibility option; use layer_options.dit.resident_layers in new configurations.
Default: 0
--host-weight-runtime-mode¶
Possible choices: disabled, preferred, required
Host Weight Runtime policy for eligible no-AllGather DLO: disabled does not consult HWR; preferred restores an exact hit or canonically loads and publishes on a miss; required restores an exact hit or fails startup. Populate a required store with preferred first.
Default: disabled
--host-weight-runtime-root¶
Writable node-local Host Weight Runtime store shared by workers in one storage domain. Required for preferred and required; use the same persistent path for population and serving.
Default: None
--dlo-host-registration-limit-gib¶
Optional per-worker GiB ceiling for registering an HWR mmap for direct H2D. Zero applies no additional ceiling. Eligible no-AllGather HWR hits attempt registration under the existing pinned-memory policy and fall back to bounded staging when unavailable.
Default: 0.0
--boundary-ratio¶
Boundary split ratio for low/high DiT in video models (e.g., 0.875 for Wan2.2).
Default: None
--flow-shift¶
Scheduler flow_shift for video models (e.g., 5.0 for 720p, 12.0 for 480p).
Default: None
--diffusion-kv-cache-dtype¶
Diffusion Q/K/V precision: fp8, mxfp8, mxfp4, or float (NPU). Separate from vLLM --kv-cache-dtype. Use --diffusion-attention-config for per-role fallback.
Default: None
--diffusion-kv-cache-skip-steps¶
Diffusion KV-cache quantization skip-step selector, e.g. '0-9,20,25-30'.
Default: None
--diffusion-kv-cache-skip-layers¶
Diffusion KV-cache quantization skip-layer selector, e.g. '0,1,4-8'.
Default: None
--cfg-parallel-size¶
Number of devices used to execute diffusion guidance passes in parallel. Equivalent to setting DiffusionParallelConfig.cfg_parallel_size.
Default: 1
--vae-patch-parallel-size¶
VAE Patch Parallelism degree for diffusion models. Distributes VAE decode workload across multiple ranks by splitting the latent spatially. Equivalent to setting DiffusionParallelConfig.vae_patch_parallel_size.
Default: 1
--text-encoder-tp-size¶
Tensor-parallel degree for the diffusion text encoder. Shards the encoder across the first N DiT ranks. Equivalent to setting DiffusionParallelConfig.text_encoder_tp_size.
Default: None
--vae-parallel-mode¶
Possible choices: tile, spatial_shard_height, spatial_shard_width
VAE parallel decode strategy for diffusion models. 'tile' (default) uses patch/tile parallel decode; 'spatial_shard_height'/'spatial_shard_width' use spatially-sharded decode that splits decoder feature maps along height/width and exchanges halo regions. The 'spatial_shard_*' modes require vae_patch_parallel_size to match the DiT group size. Equivalent to setting DiffusionParallelConfig.vae_parallel_mode.
Default: tile
--default-sampling-params¶
Json str for Default sampling parameters, Structure: {"
Default: None
--max-generated-image-size¶
Maximum generated image size in pixels (height * width).
Default: 33177600
--diffusion-streaming-output¶
Enable chunked streaming output for diffusion (mainly video generation) models that support it.
Default: False
--tts-max-instructions-length¶
Maximum length for TTS voice style instructions (overrides the pipeline default, default: 500).
Default: None
--no-guardrails¶
Disable Cosmos3 text/video safety guardrails for this server.
Default: False
--robot-openpi-idle-timeout¶
Seconds the /v1/realtime/robot/openpi endpoint waits for the next request before closing an idle WebSocket (default: 30). Set to 0 to disable the timeout.
Default: 30.0
--enable-diffusion-pipeline-profiler¶
Enable diffusion pipeline profiler to display stage durations.
Default: False
--enable-ar-profiler¶
Enable AR stage profiler to include AR stage timing in stage_durations.
Default: False
--enable-orch-monitor¶
Enable orchestrator window monitor and write a JSON log at shutdown.
Default: False
--auxiliary-text-encoder¶
Auxiliary text encoder parameters model name or path (especially for Hidream-l1-full).
Default: None
Async video jobs use process-local metadata and task handles. With multiple API workers, creating, listing, retrieving, downloading, and deleting these jobs returns HTTP 409. This restriction is independent of the current diffusion stage launch restriction: supporting multi-worker diffusion also requires a shared job lifecycle. Missing frontend topology state returns HTTP 503 for operations that require process-local state protection.