Diffusion Advanced Features¶
Table of Contents¶
Overview¶
vLLM-Omni supports various advanced features for diffusion models:
- Acceleration: cache methods, parallelism methods, startup optimizations
- Memory optimization: cpu offloading, quantization
- Extensions: LoRA inference, frame interpolation
- Execution modes: step execution
Supported Features¶
Acceleration¶
Lossy Acceleration¶
Cache methods trade minimal quality for significant speedup. Quality loss is typically imperceptible with proper tuning.
| Method | Description | Best For |
|---|---|---|
| TeaCache | Adaptive caching using modulated inputs | Quick setup, balanced quality/speed on single GPU |
| Cache-DiT | Multiple caching techniques: DBCache, TaylorSeer, SCM | Fine-grained control, tunable quality-speed tradeoff |
Diffusion KV Prefix Caching¶
KV prefix caching reuses stable context KV across requests, rather than approximating denoise steps. Enable diffusion_kv_mode: paged_scheduler and enable_prefix_caching: true on the HunyuanImage3 standalone DiT stage. This is separate from the AR stage-output cache, TeaCache and Cache-DiT.
| Model / combination | Scope |
|---|---|
| HunyuanImage3 standalone DiT | Stable text/reference-image prefix reuse; dynamic image tokens are excluded |
| TP4/SP1 + EP, CFGP1 | Validated reference-image prefix hits |
| TP2/SP2 (Ulysses) + EP, CFGP1 | Validated reference-image prefix hits; CFG guidance does not require CFG parallelism |
| Other SP/TP combinations, CFGP>1 | Prefix-hit E2E accuracy not yet validated |
| AR-imported KV / cross-stage missing-page transfer | Not covered by local prefix caching |
| Sleep mode | Rejected with prefix caching; discarded KV pages would leave stale cache hits |
| TeaCache / Cache-DiT, offload, quantization | Not validated with prefix hits |
| Step execution | Not supported with HunyuanImage3 paged KV; use request-level execution |
| Other diffusion models | No prefix-cache model adapter yet |
This table describes prefix-cache scope; the general parallelism tables below do not establish prefix-hit compatibility. Floating-point kernel differences still require accuracy validation. dense_legacy remains the default.
Lossless Acceleration¶
Parallelism methods distribute computation across GPUs without quality loss (mathematically equivalent to single-GPU).
| Method | Description | Best For |
|---|---|---|
| Ulysses-SP | Sequence parallelism via all-to-all communication | High-resolution images (>1536px) or long videos with 2-8 GPUs |
| Ring-Attention | Sequence parallelism via ring-based communication | Videos, very long sequences, memory-constrained, with 2-8 GPUs |
| CFG-Parallel | Splits CFG positive/negative branches across devices | Image editing with CFG guidance (true_cfg_scale > 1) on 2 GPUs |
| Tensor Parallelism | Shards model weights across devices | Large models that don't fit in single GPU, with 2+ GPUs |
| Pipeline Parallelism | Splits the denoising transformer block-wise across sequential GPU stages | Large diffusion transformers that need lower per-GPU model memory |
| HSDP | Weight sharding via FSDP2, redistributed on-demand at runtime | Very large models (14B+) on limited VRAM, combinable with SP |
| Expert Parallelism | Shards MoE expert MLP blocks across devices | MoE diffusion models (e.g., HunyuanImage3.0) |
Startup Optimization¶
| Method | Description | Best For |
|---|---|---|
| Multi-Thread Weight Loading | Loads safetensors shards in parallel using a thread pool | All diffusion models; reduces startup from minutes to seconds |
Note: Some acceleration methods can be combined together for optimized performance. See Feature Compatibility Table and Feature Compatibility Tutorial for detailed configuration examples.
Memory Optimization¶
Memory optimization methods help reduce GPU memory usage, enabling inference on resource-constrained hardware or larger models.
| Method | Description | Best For |
|---|---|---|
| CPU Offload | Offloads model components to CPU memory | Limited VRAM, large models on consumer GPUs |
| Quantization | Reduces transformer stages from BF16 to FP8/INT8/etc. | Limited VRAM, minimal accuracy loss |
| VAE Parallelism | Distributes VAE decode work across GPUs | High-resolution generation with reduced VAE memory peak |
Extensions¶
Extension methods add specialized capabilities to diffusion models beyond standard inference.
| Method | Description | Best For |
|---|---|---|
| LoRA Inference | Enables inference with Low-Rank Adaptation (LoRA) adapters weights | Reinforcement learning extensions |
| Frame Interpolation | Inserts intermediate video frames after generation for smoother motion | Video generation pipelines that need higher temporal smoothness |
Execution Modes¶
Execution modes control how the diffusion pipeline processes requests and denoise steps.
| Method | Description | Best For |
|---|---|---|
| Diffusion Execution Modes | Configures serial requests, request batching, step execution, continuous batching, and streaming output | Matching latency, throughput, cancellation, and output-delivery requirements |
Note: Request-level batching and step execution are capability-based. Consult the execution guide and the selected pipeline's documentation for current support.
Quantization Methods¶
| Method | Configuration | Description | Best For |
|---|---|---|---|
| FP8 | quantization="fp8" | FP8 W8A8 on validated transformer stages | Memory reduction, inference speedup |
| INT8 | quantization="int8" | INT8 W8A8 on validated transformer stages | Memory reduction, broad GPU compatibility |
| GGUF | quantization="gguf" | Native GGUF transformer-only weights (Q4, Q8, etc.) | Memory reduction on consumer GPUs |
Supported Models¶
The following tables show which models support each feature:
- 🔀SP (Ulysses & Ring): Includes both Ulysses-SP and Ring-Attention methods
- ✅ = Fully supported
- ✅* = Supported with the constraint listed below the table
- ❌ = Not supported
- ❓ = Not verified; not recommended
Notes:
- CPU Offload has three strategies: model-level, layerwise, and distributed layerwise. The tables below show layerwise support only. Split models like Cosmos3 swap their reasoner/generator components for model-level offload; see the CPU Offload Guide.
- The 💾Quantization column is collapsed for readability. See Quantization for per-method and per-model support details.
ImageGen¶
| Model | ⚡TeaCache | ⚡Cache-DiT | 🔀SP (Ulysses & Ring) | 🔀CFG-Parallel | 🔀Tensor-Parallel | 🔀Pipeline-Parallel | 🔀HSDP | 💾CPU Offload (Layerwise) | 💾VAE-Patch-Parallel | 💾Quantization | 🔄Step Execution |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Bagel | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ |
| Boogu-Image | ❌ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| FLUX.1-dev | ✅ | ✅ | ❌ | ✅ | ✅ | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ |
| FLUX.1-schnell | ❌ | ✅ | ❌ | ✅ | ✅ | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ |
| FLUX.2-klein | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ |
| FLUX.1-Kontext-dev | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ |
| FLUX.2-dev | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ | ❌ | ❌ |
| GLM-Image | ❌ | ❌ | ❌ | ✅ | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ |
| Hidream-I1-Full | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| HiDream-O1-Image | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| HunyuanImage3 | ❌ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ✅ | ✅* |
| Krea 2 | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| LongCat-Image | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| LongCat-Image-Edit | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| MammothModa2(T2I) | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Nextstep_1(T2I) | ❓ | ❓ | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| OmniGen2 | ❌ | ✅ | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Ovis-Image | ❌ | ✅ | ❌ | ✅ | ❌ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| Qwen-Image | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ✅ | ✅ |
| Qwen-Image-2512 | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ✅ | ✅ |
| Qwen-Image-Edit | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| Qwen-Image-Edit-2509 | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ (decode) | ✅ | ❌ | ❌ |
| Qwen-Image-Layered | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| SenseNova-U1 / U1.5 | ❌ | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| Stable-Diffusion-XL | ❌ | ❌ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| Stable-Diffusion3.5 | ❌ | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ✅ | ✅ (decode) | ❌ | ❌ |
| Z-Image | ✅ | ✅ | ✅ | ❓ | ✅ (TP=2 only) | ❌ | ✅ | ❌ | ✅ (decode) | ✅ | ❌ |
| ERNIE-Image | ❌ | ✅ | ✅ | ❓ | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ❌ |
| Cosmos3 | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ✅ | ❌ |
Notes:
- Nextstep_1(T2I) does not support cache acceleration methods such as TeaCache or Cache-DiT.
Tongyi-MAI/Z-Image-Turbois a distilled model with minimal NFEs; CFG-Parallel is not necessary.- Cosmos3 T2I uses
Cosmos3OmniDiffusersPipelinewithmodalities=["image"]. Model-level CPU offload swaps the nested UND reasoner and GEN generator pathways; layerwise offload remains available for blockwise GEN/UND offload.- Krea 2 currently supports single-GPU inference plus LoRA, Cache-DiT, HSDP, CPU/layerwise offload, and VAE-patch-parallel (decode). TP/SP/CFG-Parallel are not yet wired. The few-step distilled (Turbo) checkpoint uses
is_distilled=true(fixed timestep shiftmu=1.15); generate at 2048x2048 by default withnum_inference_steps≈8andguidance_scale=0. The Raw checkpoint uses 1024x1024,num_inference_steps=28, andguidance_scale=4.5.- HunyuanImage3 supports step execution. Multi-request step execution requires
TORCH_SDPA; see Diffusion Execution Modes.- BAGEL step execution supports image generation with
bagel.yaml,bagel_think.yaml, andbagel_single_stage.yaml; two-stage Thinker execution and explicit single-stage text output remain on their existing complete-request paths. Image requests requirenum_inference_steps >= 2. BAGEL step execution cannot currently be combined with sequence parallelism or a diffusion cache backend; see Diffusion Execution Modes.- MammothModa2 runs its DiT stage on the diffusion runner (
StageExecutionType.DIFFUSION); Cache-DiT is enabled through the standard diffusion-stage knobs on the stage entry of the deploy YAML (cache_backend: cache_dit, plus optionalcache_config/enable_cache_dit_summary). The runner installs the backend at startup and the pipeline adopts it per request. Only the repeated main-layer stack is cached; requests withtext_guidance_scale = 1.0bypass the cache hooks.
VideoGen¶
| Model | ⚡TeaCache | ⚡Cache-DiT | 🔀SP (Ulysses & Ring) | 🔀CFG-Parallel | 🔀Tensor-Parallel | Pipeline-Parallel | 🔀HSDP | 💾CPU Offload (Layerwise) | 💾VAE-Patch-Parallel | 💾Quantization | 🔄Step Execution |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Wan2.2 | ❌ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ (encode/decode) | ❌ | ❌ |
| Wan2.2-S2V | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (encode/decode) | ❌ | ❌ |
| Wan2.1-VACE | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| LTX-2 | ❌ | ✅ | ✅ (Ulysses only) | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| LTX-2.3 | ❌ | ✅ | ✅ (Ulysses only) | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (decode) | ❌ | ❌ |
| LTX-2.5 | ❌ | ❓ (one-stage) | ✅* (Ulysses only) | ✅ (Full only) | ❓ | ❌ | ✅ | ✅ | ✅ (decode) | ❓ (FP8) | ❌ |
| SANA-Video-2B T2V I2V | ❌ | ❌ | ✅ (--usp, frame-sharded) | ✅ | ✅ (TP=2 only) | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Helios | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ❌ | ❌ | ✅* |
| HunyuanVideo-1.5 T2V I2V | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (encode/decode) | ✅ | ❌ |
| Cosmos3 | ❌ | ✅ | ✅ | ✅ | ✅ | ❌ | ✅ | ✅ | ✅ (encode/decode) | ✅ | ❌ |
| LongCat-Video-Avatar-1.5 | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| MiniMax-H3 | ✅ (FL2VA) | ✅ | ✅ | ❌ | ✅ (DiT/TE) | ❌ | ✅ | ✅ | ✅ (tile) | ✅ (DiT) | ❌ |
| MAGI-2 Preview | ❌ | ✅ | ✅ (Ulysses) | ✅ (2-way) | ✅ | ❌ | ✅ | ✅ (1-GPU/LW; DLO DP-AG/SP no-AG) | ✅ (tile) | ❌ | ❌ |
| SANA-WM | ❌ | ❌ | ❌5 | ✅ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
Notes: 5. SANA-WM cannot support sequence parallelism: its bidirectional gated delta recurrence carries state across frames, so a rank cannot denoise a slice of the token sequence in isolation. Doing so would need a distributed scan or an all-gather before every GDN block. The remaining ❌ columns are simply unvalidated on this model, not known-broken.
Step execution note: Helios supports single-request step execution only; use
max_num_seqs=1. Frame Interpolation Support
- Supported: Wan2.2 text-to-video, image-to-video, and TI2V pipelines; SANA-Video-2B native text-to-video and image-to-video pipelines
- Not supported: Wan2.1-VACE, LTX-2, LTX-2.3, LTX-2.5, Helios, HunyuanVideo-1.5, DreamID-Omni, SANA-WM; SANA-Video-2B Diffusers-adapter text-to-video and image-to-video pipelines
AudioGen¶
| Model | ⚡TeaCache | ⚡Cache-DiT | 🔀SP (Ulysses & Ring) | 🔀CFG-Parallel | 🔀Tensor-Parallel | 🔀Pipeline-Parallel | 🔀HSDP | 💾CPU Offload (Layerwise) | 💾VAE-Patch-Parallel | 💾Quantization | 🔄Step Execution |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Stable-Audio-Open | ✅ | ❌ | ❓ | ❓ | ❌ | ❌ | ✅ | ✅ | ❌ | ✅ | ❌ |
Feature Compatibility¶
Legend:
- ✅: Functionality is supported
- ❌: No support plan
- ❓: Not verified yet and Not Recommended
| ⚡TeaCache | ⚡Cache-DiT | 🔀Ulysses-SP | 🔀Ring-Attn | 🔀CFG-Parallel | 🔀Tensor Parallel | 🔀HSDP | 🔀Expert Parallel | 💾CPU Offloading (Layerwise) | 💾CPU Offloading (Module-wise) | 💾VAE Patch Parallel | 💾FP8 Quant | 🔧LoRA Inference | 🔄Step Execution | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ⚡TeaCache | ||||||||||||||
| ⚡Cache-DiT | ❌ | |||||||||||||
| 🔀Ulysses-SP | ✅ | ✅ | ||||||||||||
| 🔀Ring-Attn | ✅ | ✅ | ✅ | |||||||||||
| 🔀CFG-Parallel | ✅ | ✅ | ✅ | ✅ | ||||||||||
| 🔀Tensor Parallel | ✅ | ✅ | ✅ | ✅ | ✅ | |||||||||
| 🔀HSDP | ❓ | ❓ | ✅ | ❓ | ✅ | ❌ | ||||||||
| 🔀Expert Parallel | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | |||||||
| 💾CPU Offloading (Layerwise) | ✅ | ✅ | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ | ||||||
| 💾CPU Offloading (Module-wise) | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❓ | ❓ | ❌ | |||||
| 💾VAE Patch Parallel | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ | ||||
| 💾FP8 Quant | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ❓ | ❓ | ✅ | ✅ | ✅ | |||
| 🔧LoRA Inference | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ❓ | ||
| 🔄Step Execution | ❌ | ❌ | ✅ | ✅ | ✅ | ✅ | ❓ | ❓ | ✅ | ❓ | ✅ | ✅ | ✅ |
Info
- HSDP can be combined with Ulysses-SP or CFG-Parallel. Tensor Parallel and HSDP are not compatible; other HSDP combinations in the table remain unverified.
- TeaCache and Cache-DiT are not compatible.
- CPU Offloading (Layerwise) and CPU Offloading (Module-wise) are not compatible.
- The CPU Offloading (Layerwise) row describes local layerwise offload. Multi-device Distributed Layerwise Offload has a separate topology and compatibility matrix in the Distributed Layerwise Offloading guide.
- The compatibility matrix uses FP8 as the representative quantization method.
- Step Execution is not compatible with any diffusion cache backend. LoRA is supported, but each scheduled batch must use a single adapter (requests with different
lora_requestorlora_scaleare kept in separate batches).
Multi-Thread Weight Loading¶
The loading guide now lives at Diffusion Startup and Loading. This heading remains so existing links to this section continue to work.
Learn More¶
The Diffusion Acceleration navigation groups the remaining guides as follows:
| Area | Guide |
|---|---|
| Compatibility | Feature Compatibility |
| CPU offloading | CPU Offloading |
| Cache acceleration | TeaCache, Cache-DiT |
| KV cache paging | Scheduler-Managed Paged KV Cache |
| Parallelism | Parallelism Overview |
| Attention | Attention Backends |
| Compilation | Regional Compilation |
| VAE decode | Wan VAE Decoder Fast Path |
| Video extension | Frame Interpolation |
| Startup | Startup and Loading |
| Adapters | LoRA |
Related cross-model and runtime features are documented separately:
- Quantization covers diffusion-only models, multi-stage omni/TTS models, and multi-stage diffusion models.
- Execution Modes and Streaming covers the diffusion runtime, including batching, step execution, and streaming output.