Skip to content

Diffusion Advanced Features

Table of Contents

Overview

vLLM-Omni supports various advanced features for diffusion models:

  • Acceleration: cache methods, parallelism methods, startup optimizations
  • Memory optimization: cpu offloading, quantization
  • Extensions: LoRA inference, frame interpolation
  • Execution modes: step execution

Supported Features

Acceleration

Lossy Acceleration

Cache methods trade minimal quality for significant speedup. Quality loss is typically imperceptible with proper tuning.

Method Description Best For
TeaCache Adaptive caching using modulated inputs Quick setup, balanced quality/speed on single GPU
Cache-DiT Multiple caching techniques: DBCache, TaylorSeer, SCM Fine-grained control, tunable quality-speed tradeoff

Diffusion KV Prefix Caching

KV prefix caching reuses stable context KV across requests, rather than approximating denoise steps. Enable diffusion_kv_mode: paged_scheduler and enable_prefix_caching: true on the HunyuanImage3 standalone DiT stage. This is separate from the AR stage-output cache, TeaCache and Cache-DiT.

Model / combination Scope
HunyuanImage3 standalone DiT Stable text/reference-image prefix reuse; dynamic image tokens are excluded
TP4/SP1 + EP, CFGP1 Validated reference-image prefix hits
TP2/SP2 (Ulysses) + EP, CFGP1 Validated reference-image prefix hits; CFG guidance does not require CFG parallelism
Other SP/TP combinations, CFGP>1 Prefix-hit E2E accuracy not yet validated
AR-imported KV / cross-stage missing-page transfer Not covered by local prefix caching
Sleep mode Rejected with prefix caching; discarded KV pages would leave stale cache hits
TeaCache / Cache-DiT, offload, quantization Not validated with prefix hits
Step execution Not supported with HunyuanImage3 paged KV; use request-level execution
Other diffusion models No prefix-cache model adapter yet

This table describes prefix-cache scope; the general parallelism tables below do not establish prefix-hit compatibility. Floating-point kernel differences still require accuracy validation. dense_legacy remains the default.

Lossless Acceleration

Parallelism methods distribute computation across GPUs without quality loss (mathematically equivalent to single-GPU).

Method Description Best For
Ulysses-SP Sequence parallelism via all-to-all communication High-resolution images (>1536px) or long videos with 2-8 GPUs
Ring-Attention Sequence parallelism via ring-based communication Videos, very long sequences, memory-constrained, with 2-8 GPUs
CFG-Parallel Splits CFG positive/negative branches across devices Image editing with CFG guidance (true_cfg_scale > 1) on 2 GPUs
Tensor Parallelism Shards model weights across devices Large models that don't fit in single GPU, with 2+ GPUs
Pipeline Parallelism Splits the denoising transformer block-wise across sequential GPU stages Large diffusion transformers that need lower per-GPU model memory
HSDP Weight sharding via FSDP2, redistributed on-demand at runtime Very large models (14B+) on limited VRAM, combinable with SP
Expert Parallelism Shards MoE expert MLP blocks across devices MoE diffusion models (e.g., HunyuanImage3.0)

Startup Optimization

Method Description Best For
Multi-Thread Weight Loading Loads safetensors shards in parallel using a thread pool All diffusion models; reduces startup from minutes to seconds

Note: Some acceleration methods can be combined together for optimized performance. See Feature Compatibility Table and Feature Compatibility Tutorial for detailed configuration examples.

Memory Optimization

Memory optimization methods help reduce GPU memory usage, enabling inference on resource-constrained hardware or larger models.

Method Description Best For
CPU Offload Offloads model components to CPU memory Limited VRAM, large models on consumer GPUs
Quantization Reduces transformer stages from BF16 to FP8/INT8/etc. Limited VRAM, minimal accuracy loss
VAE Parallelism Distributes VAE decode work across GPUs High-resolution generation with reduced VAE memory peak

Extensions

Extension methods add specialized capabilities to diffusion models beyond standard inference.

Method Description Best For
LoRA Inference Enables inference with Low-Rank Adaptation (LoRA) adapters weights Reinforcement learning extensions
Frame Interpolation Inserts intermediate video frames after generation for smoother motion Video generation pipelines that need higher temporal smoothness

Execution Modes

Execution modes control how the diffusion pipeline processes requests and denoise steps.

Method Description Best For
Diffusion Execution Modes Configures serial requests, request batching, step execution, continuous batching, and streaming output Matching latency, throughput, cancellation, and output-delivery requirements

Note: Request-level batching and step execution are capability-based. Consult the execution guide and the selected pipeline's documentation for current support.

Quantization Methods

Method Configuration Description Best For
FP8 quantization="fp8" FP8 W8A8 on validated transformer stages Memory reduction, inference speedup
INT8 quantization="int8" INT8 W8A8 on validated transformer stages Memory reduction, broad GPU compatibility
GGUF quantization="gguf" Native GGUF transformer-only weights (Q4, Q8, etc.) Memory reduction on consumer GPUs

Supported Models

The following tables show which models support each feature:

  • 🔀SP (Ulysses & Ring): Includes both Ulysses-SP and Ring-Attention methods
  • ✅ = Fully supported
  • ✅* = Supported with the constraint listed below the table
  • ❌ = Not supported
  • ❓ = Not verified; not recommended

Notes:

  1. CPU Offload has three strategies: model-level, layerwise, and distributed layerwise. The tables below show layerwise support only. Split models like Cosmos3 swap their reasoner/generator components for model-level offload; see the CPU Offload Guide.
  2. The 💾Quantization column is collapsed for readability. See Quantization for per-method and per-model support details.

ImageGen

Model ⚡TeaCache ⚡Cache-DiT 🔀SP (Ulysses & Ring) 🔀CFG-Parallel 🔀Tensor-Parallel 🔀Pipeline-Parallel 🔀HSDP 💾CPU Offload (Layerwise) 💾VAE-Patch-Parallel 💾Quantization 🔄Step Execution
Bagel ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ❌ ❌ ✅
Boogu-Image ❌ ❌ ❌ ✅ ❌ ❌ ❌ ❌ ❌ ❌ ❌
FLUX.1-dev ✅ ✅ ❌ ✅ ✅ ❌ ✅ ❌ ❌ ✅ ❌
FLUX.1-schnell ❌ ✅ ❌ ✅ ✅ ❌ ✅ ❌ ❌ ✅ ❌
FLUX.2-klein ✅ ✅ ✅ ✅ ✅ ❌ ✅ ❌ ❌ ✅ ❌
FLUX.1-Kontext-dev ❌ ✅ ❌ ❌ ✅ ❌ ✅ ❌ ❌ ❌ ❌
FLUX.2-dev ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ ❌ ❌
GLM-Image ❌ ❌ ❌ ✅ ✅ ❌ ✅ ❌ ❌ ❌ ❌
Hidream-I1-Full ❌ ❌ ❌ ❌ ✅ ❌ ❌ ❌ ❌ ❌ ❌
HiDream-O1-Image ❌ ✅ ❌ ❌ ✅ ❌ ❌ ❌ ❌ ❌ ❌
HunyuanImage3 ❌ ✅ ❌ ❌ ✅ ❌ ❌ ❌ ❌ ✅ ✅*
Krea 2 ❌ ✅ ❌ ❌ ❌ ❌ ✅ ✅ ✅ (decode) ❌ ❌
LongCat-Image ✅ ✅ ✅ ✅ ✅ ❌ ❌ ✅ ❌ ❌ ❌
LongCat-Image-Edit ✅ ✅ ✅ ✅ ✅ ❌ ❌ ✅ ❌ ❌ ❌
MammothModa2(T2I) ❌ ✅ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌
Nextstep_1(T2I) ❓ ❓ ❌ ✅ ✅ ❌ ❌ ✅ ❌ ❌ ❌
OmniGen2 ❌ ✅ ✅ ❌ ✅ ❌ ❌ ❌ ❌ ❌ ❌
Ovis-Image ❌ ✅ ❌ ✅ ❌ ❌ ❌ ✅ ❌ ❌ ❌
Qwen-Image ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ✅ ✅
Qwen-Image-2512 ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ✅ ✅
Qwen-Image-Edit ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
Qwen-Image-Edit-2509 ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ (decode) ✅ ❌ ❌
Qwen-Image-Layered ✅ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
SenseNova-U1 / U1.5 ❌ ✅ ❌ ✅ ✅ ❌ ❌ ✅ ❌ ❌ ❌
Stable-Diffusion-XL ❌ ❌ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
Stable-Diffusion3.5 ❌ ✅ ❌ ✅ ✅ ❌ ❌ ✅ ✅ (decode) ❌ ❌
Z-Image ✅ ✅ ✅ ❓ ✅ (TP=2 only) ❌ ✅ ❌ ✅ (decode) ✅ ❌
ERNIE-Image ❌ ✅ ✅ ❓ ✅ ❌ ✅ ✅ ❌ ❌ ❌
Cosmos3 ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ✅ ❌

Notes:

  1. Nextstep_1(T2I) does not support cache acceleration methods such as TeaCache or Cache-DiT.
  2. Tongyi-MAI/Z-Image-Turbo is a distilled model with minimal NFEs; CFG-Parallel is not necessary.
  3. Cosmos3 T2I uses Cosmos3OmniDiffusersPipeline with modalities=["image"]. Model-level CPU offload swaps the nested UND reasoner and GEN generator pathways; layerwise offload remains available for blockwise GEN/UND offload.
  4. Krea 2 currently supports single-GPU inference plus LoRA, Cache-DiT, HSDP, CPU/layerwise offload, and VAE-patch-parallel (decode). TP/SP/CFG-Parallel are not yet wired. The few-step distilled (Turbo) checkpoint uses is_distilled=true (fixed timestep shift mu=1.15); generate at 2048x2048 by default with num_inference_steps≈8 and guidance_scale=0. The Raw checkpoint uses 1024x1024, num_inference_steps=28, and guidance_scale=4.5.
  5. HunyuanImage3 supports step execution. Multi-request step execution requires TORCH_SDPA; see Diffusion Execution Modes.
  6. BAGEL step execution supports image generation with bagel.yaml, bagel_think.yaml, and bagel_single_stage.yaml; two-stage Thinker execution and explicit single-stage text output remain on their existing complete-request paths. Image requests require num_inference_steps >= 2. BAGEL step execution cannot currently be combined with sequence parallelism or a diffusion cache backend; see Diffusion Execution Modes.
  7. MammothModa2 runs its DiT stage on the diffusion runner (StageExecutionType.DIFFUSION); Cache-DiT is enabled through the standard diffusion-stage knobs on the stage entry of the deploy YAML (cache_backend: cache_dit, plus optional cache_config / enable_cache_dit_summary). The runner installs the backend at startup and the pipeline adopts it per request. Only the repeated main-layer stack is cached; requests with text_guidance_scale = 1.0 bypass the cache hooks.

VideoGen

Model ⚡TeaCache ⚡Cache-DiT 🔀SP (Ulysses & Ring) 🔀CFG-Parallel 🔀Tensor-Parallel Pipeline-Parallel 🔀HSDP 💾CPU Offload (Layerwise) 💾VAE-Patch-Parallel 💾Quantization 🔄Step Execution
Wan2.2 ❌ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ (encode/decode) ❌ ❌
Wan2.2-S2V ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (encode/decode) ❌ ❌
Wan2.1-VACE ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
LTX-2 ❌ ✅ ✅ (Ulysses only) ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
LTX-2.3 ❌ ✅ ✅ (Ulysses only) ✅ ✅ ❌ ✅ ✅ ✅ (decode) ❌ ❌
LTX-2.5 ❌ ❓ (one-stage) ✅* (Ulysses only) ✅ (Full only) ❓ ❌ ✅ ✅ ✅ (decode) ❓ (FP8) ❌
SANA-Video-2B T2V I2V ❌ ❌ ✅ (--usp, frame-sharded) ✅ ✅ (TP=2 only) ❌ ❌ ❌ ❌ ❌ ❌
Helios ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ❌ ❌ ✅*
HunyuanVideo-1.5 T2V I2V ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (encode/decode) ✅ ❌
Cosmos3 ❌ ✅ ✅ ✅ ✅ ❌ ✅ ✅ ✅ (encode/decode) ✅ ❌
LongCat-Video-Avatar-1.5 ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌ ❌
MiniMax-H3 ✅ (FL2VA) ✅ ✅ ❌ ✅ (DiT/TE) ❌ ✅ ✅ ✅ (tile) ✅ (DiT) ❌
MAGI-2 Preview ❌ ✅ ✅ (Ulysses) ✅ (2-way) ✅ ❌ ✅ ✅ (1-GPU/LW; DLO DP-AG/SP no-AG) ✅ (tile) ❌ ❌
SANA-WM ❌ ❌ ❌5 ✅ ✅ ❌ ❌ ❌ ❌ ❌ ❌

Notes: 5. SANA-WM cannot support sequence parallelism: its bidirectional gated delta recurrence carries state across frames, so a rank cannot denoise a slice of the token sequence in isolation. Doing so would need a distributed scan or an all-gather before every GDN block. The remaining ❌ columns are simply unvalidated on this model, not known-broken.

Step execution note: Helios supports single-request step execution only; use max_num_seqs=1. Frame Interpolation Support

  • Supported: Wan2.2 text-to-video, image-to-video, and TI2V pipelines; SANA-Video-2B native text-to-video and image-to-video pipelines
  • Not supported: Wan2.1-VACE, LTX-2, LTX-2.3, LTX-2.5, Helios, HunyuanVideo-1.5, DreamID-Omni, SANA-WM; SANA-Video-2B Diffusers-adapter text-to-video and image-to-video pipelines

AudioGen

Model ⚡TeaCache ⚡Cache-DiT 🔀SP (Ulysses & Ring) 🔀CFG-Parallel 🔀Tensor-Parallel 🔀Pipeline-Parallel 🔀HSDP 💾CPU Offload (Layerwise) 💾VAE-Patch-Parallel 💾Quantization 🔄Step Execution
Stable-Audio-Open ✅ ❌ ❓ ❓ ❌ ❌ ✅ ✅ ❌ ✅ ❌

Feature Compatibility

Legend:

  • ✅: Functionality is supported
  • ❌: No support plan
  • ❓: Not verified yet and Not Recommended
⚡TeaCache ⚡Cache-DiT 🔀Ulysses-SP 🔀Ring-Attn 🔀CFG-Parallel 🔀Tensor Parallel 🔀HSDP 🔀Expert Parallel 💾CPU Offloading (Layerwise) 💾CPU Offloading (Module-wise) 💾VAE Patch Parallel 💾FP8 Quant 🔧LoRA Inference 🔄Step Execution
⚡TeaCache
⚡Cache-DiT ❌
🔀Ulysses-SP ✅ ✅
🔀Ring-Attn ✅ ✅ ✅
🔀CFG-Parallel ✅ ✅ ✅ ✅
🔀Tensor Parallel ✅ ✅ ✅ ✅ ✅
🔀HSDP ❓ ❓ ✅ ❓ ✅ ❌
🔀Expert Parallel ❓ ❓ ❓ ❓ ❓ ❓ ❓
💾CPU Offloading (Layerwise) ✅ ✅ ❌ ❌ ❌ ❌ ❌ ❌
💾CPU Offloading (Module-wise) ✅ ✅ ✅ ✅ ✅ ✅ ❓ ❓ ❌
💾VAE Patch Parallel ✅ ✅ ✅ ✅ ✅ ✅ ✅ ✅ ❌ ❌
💾FP8 Quant ✅ ✅ ✅ ✅ ✅ ✅ ❓ ❓ ✅ ✅ ✅
🔧LoRA Inference ❓ ❓ ❓ ❓ ❓ ❓ ❓ ❓ ❓ ❓ ❓ ❓
🔄Step Execution ❌ ❌ ✅ ✅ ✅ ✅ ❓ ❓ ✅ ❓ ✅ ✅ ✅

Info

  1. HSDP can be combined with Ulysses-SP or CFG-Parallel. Tensor Parallel and HSDP are not compatible; other HSDP combinations in the table remain unverified.
  2. TeaCache and Cache-DiT are not compatible.
  3. CPU Offloading (Layerwise) and CPU Offloading (Module-wise) are not compatible.
  4. The CPU Offloading (Layerwise) row describes local layerwise offload. Multi-device Distributed Layerwise Offload has a separate topology and compatibility matrix in the Distributed Layerwise Offloading guide.
  5. The compatibility matrix uses FP8 as the representative quantization method.
  6. Step Execution is not compatible with any diffusion cache backend. LoRA is supported, but each scheduled batch must use a single adapter (requests with different lora_request or lora_scale are kept in separate batches).

Multi-Thread Weight Loading

The loading guide now lives at Diffusion Startup and Loading. This heading remains so existing links to this section continue to work.

Learn More

The Diffusion Acceleration navigation groups the remaining guides as follows:

Area Guide
Compatibility Feature Compatibility
CPU offloading CPU Offloading
Cache acceleration TeaCache, Cache-DiT
KV cache paging Scheduler-Managed Paged KV Cache
Parallelism Parallelism Overview
Attention Attention Backends
Compilation Regional Compilation
VAE decode Wan VAE Decoder Fast Path
Video extension Frame Interpolation
Startup Startup and Loading
Adapters LoRA

Related cross-model and runtime features are documented separately:

  • Quantization covers diffusion-only models, multi-stage omni/TTS models, and multi-stage diffusion models.
  • Execution Modes and Streaming covers the diffusion runtime, including batching, step execution, and streaming output.