Skip to content

Quickstart

This guide will help you quickly get started with vLLM-Omni to perform:

  • Offline batched inference
  • Online serving using OpenAI-compatible server

Prerequisites

  • OS: Linux
  • Python: 3.12

Installation

For installation on GPU from source:

uv venv --python 3.12 --seed
source .venv/bin/activate

# On CUDA
uv pip install vllm==0.30.0 --torch-backend=auto

# On ROCm
uv pip install vllm==0.30.0+rocm723 --extra-index-url https://wheels.vllm.ai/rocm/0.30.0/rocm723

git clone https://github.com/vllm-project/vllm-omni.git
cd vllm-omni
uv pip install -e .

For additional installation methods — please see the installation guide.

Note

It is important to install the same major & minor version of vLLM and vLLM Omni, otherwise things may not work as expected. If the versions are misaligned, you will see a warning when you import vLLM Omni.

If the vllm command does not handle --omni correctly, verify that the active environment contains the matching vLLM release and vLLM-Omni installation. vLLM owns the CLI entrypoint and delegates --omni requests to vLLM-Omni.

Offline Inference

Text-to-image generation quickstart with vLLM-Omni:

from vllm_omni.entrypoints.omni import Omni

if __name__ == "__main__":
    omni = Omni(model="Tongyi-MAI/Z-Image-Turbo")
    prompt = "a cup of coffee on the table"
    outputs = omni.generate(prompt)
    images = outputs[0].images
    images[0].save("coffee.png")

You can pass a list of prompts and wait for the independent requests to finish, as shown below.

Info

For diffusion pipelines, each prompt becomes a separate logical request. The runtime may automatically batch compatible in-flight requests through the scheduler and runner.

from vllm_omni.entrypoints.omni import Omni

if __name__ == "__main__":
    omni = Omni(
        model="Tongyi-MAI/Z-Image-Turbo",
        # deploy_config="./deploy-config.yaml",  # Optional deploy override
    )
    prompts = [
        "a cup of coffee on a table",
        "a toy dinosaur on a sandy beach",
        "a fox waking up in bed and yawning",
    ]
    omni_outputs = omni.generate(prompts)
    for i_prompt, prompt_output in enumerate(omni_outputs):
        this_images = prompt_output.images
        for i_image, image in enumerate(this_images):
            image.save(f"p{i_prompt}-img{i_image}.jpg")
            print("saved to", f"p{i_prompt}-img{i_image}.jpg")
            # saved to p0-img0.jpg
            # saved to p1-img0.jpg
            # saved to p2-img0.jpg

Info

For diffusion request batching, step execution, and streaming controls, see Diffusion Execution Modes.

For more usages, please refer to offline inference

Online Serving with the API Server

Text-to-image generation quickstart with vLLM-Omni:

vllm serve Tongyi-MAI/Z-Image-Turbo --omni --port 8091
curl -s http://localhost:8091/v1/images/generations \
  -H "Content-Type: application/json" \
  -d '{
    "prompt": "a cup of coffee on the table",
    "size": "1024x1024",
    "response_format": "b64_json",
    "seed": 42
  }' | jq -r '.data[0].b64_json' | base64 -d > coffee.png

See the API Server guide to choose an endpoint. For model-specific details, refer to the text-to-image online serving example.