Skip to content

vllm-omni serve

Multiple API frontends

For local EngineCore pipelines with one or more stages, --api-server-count starts multiple API frontend processes that share one parent-owned set of stage engines:

vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
    --api-server-count 2 \
    --port 8091

This mode is intended for workloads whose request parsing, multimodal input processing, or response encoding can bottleneck a single frontend. It currently requires local multi-process EngineCore stages with one local process group per replica. A stage may still use multiple num_replicas. It cannot be combined with diffusion, headless or remote stages, intra-stage data parallelism, Ray, fault tolerance, elastic expert parallelism, sleep mode, or runtime LoRA updating. Runtime voice upload and deletion are also disabled because those mutations are process-local; built-in voices, inline reference audio, and voices restored at startup remain available. The /v1/omni/sleep and /v1/omni/wakeup control routes return HTTP 409 for the same reason: their bookkeeping is process-local while stage engines are shared.

A single-stage model pipeline is supported through the normal launch command; this is distinct from the distributed --stage-id / --headless launch mode. Every stage in the pipeline must be an EngineCore stage. Mixed AR-to-DiT pipelines, including MiniMax-H3, remain unsupported even when their first stage is an EngineCore stage. Diffusion support requires per-frontend channels and shared-engine lifecycle handling in the diffusion backend.

Workloads and performance validation

Each API process owns its frontend processor, Orchestrator, request state, and per-replica EngineCore channels. Requests return through the channel of the originating API process; model weights and GPU stage engines are shared. This distributes frontend and orchestration work across CPU processes.

Workload Potential benefit Limitation
Concurrent cold image/video inputs on supported EngineCore pipelines Parallel frontend processing when one API process is CPU-limited A single cold request and GPU computation may be unchanged
Many replicas producing frequent text/audio outputs Less API/Orchestrator work and GIL contention per process Routing remains serial within each Orchestrator
Warm inputs or low-concurrency, GPU-limited generation Additional frontend capacity More API processes may add overhead without improving throughput

Here, a cold input means media or preprocessing caches miss, not that model weights must be loaded or GPU kernels compiled. Each frontend has its own processor/cache state, so cache duplication, CPU thread contention, and memory usage also matter. Multi-API serving does not itself change the output queue, move response encoding to an executor, or change the model's numerical scheduler.

Compare --api-server-count 1, 2, and higher counts with the same hardware, model/stage configuration, media, sampling parameters, output lengths, and offered load. Keep cold-cache and warmed-cache runs separate and report repeated measurements. Record CPU and request counts for every API worker, input and output queue delay, GPU utilization, first-output latency, and end-to-end latency. Confirm requests reach every worker: connection reuse can concentrate traffic on one process. Queue-only or single-API experiments do not establish a multi-API throughput gain. The multi-replica workload in #4680 remains a relevant validation target.

Stage-based CLI quickstart

The stage-based CLI is designed for deployments that require launching each pipeline stage in an isolated process (e.g., across separate operating system processes, distinct GPUs, or distributed hosts).

  • For migrated models that utilize the bundled deployment YAML configurations located in vllm_omni/deploy/, the --deploy-config flag is only required to override the default configuration. By default, executing vllm serve MODEL --omni ... automatically loads the bundled deployment configuration.

Example: Initializing Stage 0 (Orchestrator and API Server): The commands below show a common device mapping where Stage 0 uses GPU 0 and worker stages use GPU 1 via CUDA_VISIBLE_DEVICES.

CUDA_VISIBLE_DEVICES=0 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
    --port 8091 \
    --stage-id 0 \
    --omni-master-address 127.0.0.1 \
    --omni-master-port 26000

Example: Initializing a Headless Worker Stage (Stage 1):

CUDA_VISIBLE_DEVICES=1 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
    --stage-id 1 \
    --headless \
    --omni-master-address 127.0.0.1 \
    --omni-master-port 26000

When utilizing a custom deployment YAML, append --deploy-config /path/to/override.yaml to each command execution.

In the standard execution paradigm, the --stage-overrides argument is utilized to apply stage-specific configurations from a single CLI command. However, under the stage-based CLI paradigm, where each process strictly encapsulates a single stage, it is recommended to specify tuning parameters directly via discrete command-line flags for the respective stage, rather than constructing a composite --stage-overrides JSON string.

For example, as an alternative to the following composite configuration:

vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni --port 8091 \
    --stage-overrides '{"1": {"gpu_memory_utilization": 0.5}}'

the stage-based CLI permits the direct initialization of Stage 1 with explicit parameters:

CUDA_VISIBLE_DEVICES=1 vllm serve Qwen/Qwen3-Omni-30B-A3B-Instruct --omni \
    --stage-id 1 \
    --headless \
    --gpu-memory-utilization 0.5 \
    --omni-master-address 127.0.0.1 \
    --omni-master-port 26000

JSON CLI Arguments

Arguments

OmniConfig

Configuration for vLLM-Omni multi-stage and diffusion models.

--omni

Enable vLLM-Omni mode for multi-modal and diffusion models

Default: False

--enable-sleep-mode

Enable GPU memory pool for sleep mode.

Default: False

--task-type

Model-defined startup task type. The selected model validates supported values; for example, TTS models accept CustomVoice, VoiceDesign, or Base, while diffusion models may use it to select task-specific weights. If omitted, the model default is used.

Default: None

--forced-aligner

Enable streaming TTS word timestamps via a forced aligner. Pass the aligner model path/name, e.g. 'Qwen/Qwen3-ForcedAligner-0.6B'. Disabled when omitted.

Default: None

--forced-aligner-config

Optional YAML file for forced aligner settings (model, runner, gpu_memory_utilization, dtype, max_model_len). The --forced-aligner flag, when set, overrides the YAML model field.

Default: None

--forced-aligner-device

Device(s) for the forced-aligner stage (e.g. '2'). Defaults to sharing an existing stage's GPU when unset.

Default: None

--deploy-config

Path to a deploy config YAML (new format with stages/engine_args).

Default: None

--strategy-config

Path to a composable-parallel strategy.yaml. Only applies to registry-based models (e.g. qwen2_5_omni): its derived parallel sizing is overlaid onto the registry-merged stages before per-stage engine args are built.

Default: None

--stage-overrides

Per-stage JSON overrides. Example: '{"0": {"gpu_memory_utilization": 0.8}, "2": {"enforce_eager": true}}'

Default: None

--async-chunk, --no-async-chunk

Override the deploy YAML's async_chunk: bool. Unset leaves the YAML value in force.

Default: None

--stage-id

Select and launch a single stage by stage_id.

Default: None

--replica-id

Deprecated and ignored — replica ids are auto-assigned by the master server. Specifying this flag prints a warning and has no effect.

Default: None

--stage-init-timeout

The timeout for initializing a single stage in seconds (default: 300)

Default: 300

--init-timeout

The timeout for initializing the stages.

Default: 600

--shm-threshold-bytes

The threshold for the shared memory size.

Default: 65536

--log-stats

Enable logging the stats.

Default: False

--log-file

The path to the log file.

Default: None

--batch-timeout

The timeout for the batch.

Default: 10

--worker-backend

Possible choices: multi_process, ray

The backend to use for stage workers.

Default: multi_process

--ray-address

The address of the Ray cluster to connect to.

Default: None

--omni-master-address, -oma

Hostname or IP address of the Omni orchestrator (master).

Default: None

--omni-master-port, -omp

Port of the Omni orchestrator (master).

Default: None

--omni-replica-address, -ora

Local bind address (this host's IP) that the headless stage advertises to the Omni master for its handshake/input/output ZMQ sockets. If unset, auto-detected via a UDP-connect routing probe against --omni-master-address. Override only when the auto-detected IP is wrong (e.g. multi-NIC host where the master is reachable on the wrong interface).

Default: None

--omni-dp-size-local

Number of stage replicas this runtime launches locally for its own --stage-id. Process-local: head and every headless invocation read their own copy; values may differ across invocations. Requires --stage-id to be set when not equal to 1.

Default: 1

--omni-lb-policy

Possible choices: random, round-robin, least-queue-length

Per-stage load-balancing policy used by the head's StagePool to route requests across UP replicas. Only consulted on the head runtime.

Default: random

--omni-heartbeat-timeout

Seconds before an unreporting replica is marked ERROR in the OmniCoordinator. Only consulted on the head runtime.

Default: 30.0

--num-gpus

Number of GPUs to use for diffusion model inference.

Default: None

--model-class-name

Override the diffusion pipeline class name (e.g. LTX2Pipeline).

Default: None

--diffusion-load-format

Possible choices: default, custom_pipeline, dummy, diffusers

How to load the diffusion pipeline: native/registry (default), custom_pipeline, dummy, or diffusers for the HF diffusers adapter.

Default: None

--lora-path

LoRA checkpoint path(s) loaded when the diffusion server starts. The distill backend accepts one file per pipeline transformer.

Default: None

--lora-backend

Possible choices: peft, distill

Diffusion LoRA loading backend. 'distill' fuses checkpoint files into the base model at server startup; 'peft' uses the adapter manager.

Default: None

--lora-scale

Scale for a startup PEFT LoRA. Distilled LoRAs are fused at their checkpoint scale.

Default: None

--diffusion-compile-granularity

Possible choices: regional, full

Compilation scope for the generic diffusion model runner. 'regional' compiles repeated blocks (default); 'full' compiles the whole transformer and is incompatible with HSDP, sequence parallelism, CPU offload, and layerwise offload.

Default: None

--diffusion-compile-dynamic, --no-diffusion-compile-dynamic

Use dynamic shapes for the selected generic diffusion compile scope. Disable for fixed-shape workloads with --no-diffusion-compile-dynamic.

Default: None

--fa-deterministic

Request FlashAttention deterministic=True on the local FLASH_ATTN dense path. Slower than the library default deterministic=False; intended for accuracy CI. Serving default remains non-deterministic.

Default: False

--diffusers-load-kwargs

JSON object passed to DiffusionPipeline.from_pretrained().It overrides corresponding parameters in the standard vLLM-Omni interface.(e.g. '{"use_safetensors": true, "variant": "fp16"}').

Default: {}

--diffusers-call-kwargs

JSON object passed to pipeline.call(). Useful for model-specific sampling parameters not covered by the vLLM-Omni interface.During request time, it is overridden by corresponding parameters in the vLLM-Omni interface.(e.g. '{"num_inference_steps": 30, "guidance_scale": 7.5}').

Default: {}

--custom-pipeline-args

JSON object passed to native/custom diffusion pipelines. Only args containing "pipeline_class" trigger custom pipeline re-initialization.

Default: None

--usp, --ulysses-degree

Ulysses Sequence Parallelism degree for diffusion models. Equivalent to setting DiffusionParallelConfig.ulysses_degree.

Default: None

--ulysses-mode

Possible choices: strict, advanced_uaa

Ulysses sequence-parallel mode for diffusion models. 'strict' keeps the original divisibility requirements; 'advanced_uaa' enables the experimental UAA path for uneven sequence/head shapes.

Default: strict

--ulysses-a2a-permute, --no-ulysses-a2a-permute

Enable fused permute-free Ulysses all-to-all over NCCL symmetric memory. Only strict Ulysses layouts are eligible. Defaults to disabled.

Default: None

--ring, --ring-degree

Ring Sequence Parallelism degree for diffusion models. Equivalent to setting DiffusionParallelConfig.ring_degree.

Default: None

--allgather-degree

AllGather-KV Sequence Parallelism degree for non-causal diffusion attention. Equivalent to setting DiffusionParallelConfig.allgather_degree.

Default: None

--diffusion-quantization-config

JSON string for diffusion quantization_config. Example: '{"method":"fp8","activation_scheme":"dynamic"}'.

Default: None

--force-cutlass-fp8

Diffusion-only runtime override for ModelOpt FP8 checkpoints: force CUTLASS FP8 linear kernels on CUDA SM89+ devices. Ignored for BF16, non-ModelOpt FP8, ROCm, and older CUDA GPUs.

Default: None

--use-hsdp

Enable HSDP (Hybrid Sharded Data Parallel) for diffusion models. Shards model weights across GPUs to reduce per-GPU memory usage.

Default: False

--hsdp-shard-size

Number of GPUs to shard weights across. -1 = auto (world_size / replicate_size).

Default: -1

--hsdp-replicate-size

Number of replica groups for HSDP. Each group holds a full sharded copy.

Default: 1

--diffusion-attention-backend

Diffusion attention backend (shorthand). Sets the default backend for all diffusion attention roles, e.g. 'FLASH_ATTN'. May be combined with --diffusion-attention-config.per_role.* overrides, but mutually exclusive with --diffusion-attention-config.default.backend.

Default: None

--fastvideo-vsa-topk

Number of key/value blocks selected per query block by FASTVIDEO_VSA.

Default: None

--diffusion-attention-config, -dac

Diffusion attention config. Accepts JSON or vLLM-style dotted flags. Examples: --diffusion-attention-config.default.backend FLASH_ATTN, --diffusion-attention-config.default.backend TRTLLM_ATTN --diffusion-attention-config.default.skip_softmax.target_sparsity 0.5, --diffusion-attention-config.per_role.cross.backend SAGE_ATTN, --diffusion-attention-config '{"default": {"backend": "FLASH_ATTN"}, "per_role": {"cross": {"backend": "SAGE_ATTN"}}}'.

Default: None

--cache-backend

Cache backend for diffusion models, options: 'tea_cache', 'cache_dit', 'mag_cache', 'sea_cache', 'step_cache'

Default: none

--cache-config

JSON string of cache configuration. TeaCache: '{"rel_l1_thresh": 0.2}'. MagCache: '{"mag_threshold": 0.24, "mag_max_skip_steps": 5, "mag_retention_ratio": 0.1}'. Calibration mode: add '"mag_calibrate": true'

Default: None

--video-output-transport

JSON object configuring video output preparation, for example '{"enable_device_postprocess": true}'.

Default: None

--enable-cache-dit-summary

Enable cache-dit summary logging after diffusion forward passes.

Default: False

--step-execution

Enable per-step diffusion execution so running requests can be aborted between denoise steps.

Default: False

--request-batch-max-wait-ms

Request-mode batch admission: max milliseconds to wait for compatible requests to accumulate before scheduling a fused forward wave. 0 disables admission (default).

Default: 0.0

--vae-use-slicing

Enable VAE slicing for memory optimization (useful for mitigating OOM issues).

Default: False

--vae-use-tiling

Enable VAE tiling for memory optimization (useful for mitigating OOM issues).

Default: False

--vae-fast-path

Possible choices: off, lossless, channels_last

Wan VAE decoder fast path. 'lossless' (default) installs bit-exact fused kernels; 'channels_last' additionally switches decoder convolutions to channels-last memory format and fuses RMSNorm+SiLU (faster, not bit-exact); 'off' keeps the reference diffusers implementation.

Default: lossless

--disable-multithread-weight-load

Disable multi-threaded safetensors loading (default: enabled with 4 threads).

Default: True

--enable-broadcast-weight-load

Enable Rank-0 shared weight broadcast across workers for HSDP (default: disabled).

Default: False

--num-weight-load-threads

Number of threads for parallel weight loading (default: 4).

Default: 4

--diffusion-offload-config

Diffusion CPU-offload config as JSON. Set mode to module or layer, list dit and/or text_encoder in components, and put layer-only tuning under layer_options. Layer settings are weight_transfer (rank-local or allgather) and resident_layers (DiT only).

Default: None

--enable-cpu-offload

Compatibility alias for model-level CPU offload. New integrations should use --diffusion-offload-config with mode=module and explicit components.

Default: False

--enable-layerwise-offload

Compatibility alias for layerwise CPU offload. New integrations should use --diffusion-offload-config with mode=layer and explicit components.

Default: False

--enable-distributed-layerwise-offload

Compatibility alias for distributed layerwise CPU offload. New integrations should use mode=layer and configure weight transfer per component.

Default: False

--dlo-use-allgather

Compatibility option; use component weight_transfer=allgather in new configurations. Use shard + AllGather for weight reconstruction (default: True). When disabled (--dlo-no-use-allgather), each rank streams the standard loader's rank-local tensors via H2D only — no additional DP sharding, no AllGather, and no concurrent-request requirement.

Default: True

--dlo-no-use-allgather

Compatibility option; use component weight_transfer=rank-local in new configurations. Disable AllGather and stream standard-loader rank-local weights independently (including existing TP shards).

Default: True

--dlo-resident-layers

Compatibility option; use layer_options.dit.resident_layers in new configurations.

Default: 0

--host-weight-runtime-mode

Possible choices: disabled, preferred, required

Host Weight Runtime policy for eligible no-AllGather DLO: disabled does not consult HWR; preferred restores an exact hit or canonically loads and publishes on a miss; required restores an exact hit or fails startup. Populate a required store with preferred first.

Default: disabled

--host-weight-runtime-root

Writable node-local Host Weight Runtime store shared by workers in one storage domain. Required for preferred and required; use the same persistent path for population and serving.

Default: None

--dlo-host-registration-limit-gib

Optional per-worker GiB ceiling for registering an HWR mmap for direct H2D. Zero applies no additional ceiling. Eligible no-AllGather HWR hits attempt registration under the existing pinned-memory policy and fall back to bounded staging when unavailable.

Default: 0.0

--boundary-ratio

Boundary split ratio for low/high DiT in video models (e.g., 0.875 for Wan2.2).

Default: None

--flow-shift

Scheduler flow_shift for video models (e.g., 5.0 for 720p, 12.0 for 480p).

Default: None

--diffusion-kv-cache-dtype

Diffusion Q/K/V precision: fp8, mxfp8, mxfp4, or float (NPU). Separate from vLLM --kv-cache-dtype. Use --diffusion-attention-config for per-role fallback.

Default: None

--diffusion-kv-cache-skip-steps

Diffusion KV-cache quantization skip-step selector, e.g. '0-9,20,25-30'.

Default: None

--diffusion-kv-cache-skip-layers

Diffusion KV-cache quantization skip-layer selector, e.g. '0,1,4-8'.

Default: None

--cfg-parallel-size

Number of devices used to execute diffusion guidance passes in parallel. Equivalent to setting DiffusionParallelConfig.cfg_parallel_size.

Default: 1

--vae-patch-parallel-size

VAE Patch Parallelism degree for diffusion models. Distributes VAE decode workload across multiple ranks by splitting the latent spatially. Equivalent to setting DiffusionParallelConfig.vae_patch_parallel_size.

Default: 1

--text-encoder-tp-size

Tensor-parallel degree for the diffusion text encoder. Shards the encoder across the first N DiT ranks. Equivalent to setting DiffusionParallelConfig.text_encoder_tp_size.

Default: None

--vae-parallel-mode

Possible choices: tile, spatial_shard_height, spatial_shard_width

VAE parallel decode strategy for diffusion models. 'tile' (default) uses patch/tile parallel decode; 'spatial_shard_height'/'spatial_shard_width' use spatially-sharded decode that splits decoder feature maps along height/width and exchanges halo regions. The 'spatial_shard_*' modes require vae_patch_parallel_size to match the DiT group size. Equivalent to setting DiffusionParallelConfig.vae_parallel_mode.

Default: tile

--default-sampling-params

Json str for Default sampling parameters, Structure: {"": {: value, ...}, ...} e.g., '{"0": {"num_inference_steps":50, "guidance_scale":1}}'. Currently only supports diffusion models.

Default: None

--max-generated-image-size

Maximum generated image size in pixels (height * width).

Default: 33177600

--diffusion-streaming-output

Enable chunked streaming output for diffusion (mainly video generation) models that support it.

Default: False

--tts-max-instructions-length

Maximum length for TTS voice style instructions (overrides the pipeline default, default: 500).

Default: None

--no-guardrails

Disable Cosmos3 text/video safety guardrails for this server.

Default: False

--robot-openpi-idle-timeout

Seconds the /v1/realtime/robot/openpi endpoint waits for the next request before closing an idle WebSocket (default: 30). Set to 0 to disable the timeout.

Default: 30.0

--enable-diffusion-pipeline-profiler

Enable diffusion pipeline profiler to display stage durations.

Default: False

--enable-ar-profiler

Enable AR stage profiler to include AR stage timing in stage_durations.

Default: False

--enable-orch-monitor

Enable orchestrator window monitor and write a JSON log at shutdown.

Default: False

--auxiliary-text-encoder

Auxiliary text encoder parameters model name or path (especially for Hidream-l1-full).

Default: None

Async video jobs use process-local metadata and task handles. With multiple API workers, creating, listing, retrieving, downloading, and deleting these jobs returns HTTP 409. This restriction is independent of the current diffusion stage launch restriction: supporting multi-worker diffusion also requires a shared job lifecycle. Missing frontend topology state returns HTTP 503 for operations that require process-local state protection.