Skip to content

vllm_omni.diffusion.models.boogu_image.image_processor

Native port of the upstream BooguImageProcessor.

Ported from the external boogu package (boogu/pipelines/image_processor.py) so the native vLLM-Omni pipeline can reproduce Boogu-Image's reference-image and VLM preprocessing without requiring pip install boogu-image.

Only the two Boogu-specific methods are ported: get_new_height_width (the max_pixels / max_side_length aware downscale-only resize) and preprocess (which routes through it). Everything else is inherited from the stock diffusers VaeImageProcessor.

BooguImageProcessor

Bases: VaeImageProcessor

VaeImageProcessor variant with Boogu-Image pixel/side-length constraints.

Resizing never upscales (the ratio is clamped to <= 1) and always aligns the target height/width to multiples of vae_scale_factor.

max_pixels instance-attribute

max_pixels = max_pixels

max_side_length instance-attribute

max_side_length = max_side_length

get_new_height_width

get_new_height_width(
    image: Image | ndarray | Tensor,
    height: int | None = None,
    width: int | None = None,
    max_pixels: int | None = None,
    max_side_length: int | None = None,
) -> tuple[int, int]

Return target (height, width) after downscale + alignment.

Faithful port of upstream BooguImageProcessor.get_new_height_width.

preprocess

preprocess(
    image: PipelineImageInput,
    height: int | None = None,
    width: int | None = None,
    max_pixels: int | None = None,
    max_side_length: int | None = None,
    resize_mode: str = "default",
    crops_coords: tuple[int, int, int, int] | None = None,
) -> Tensor

Preprocess an image into a normalized [B, C, H, W] tensor.

Faithful port of upstream BooguImageProcessor.preprocess (PixArt-style downscale). Only the PIL branch is exercised by the native pipeline, but the numpy/tensor branches are kept for parity.