Skip to content

vllm_omni.diffusion.models.boogu_image.pipeline_boogu_image

Native vLLM-Omni pipeline for Boogu-Image-0.1.

Ported from the upstream boogu package (boogu/pipelines/boogu/pipeline_boogu.py) with the following changes:

  • Diffusers DiffusionPipeline/register_modules machinery replaced by a plain nn.Module constructed from OmniDiffusionConfig (components are loaded from the checkpoint subfolders; transformer weights arrive later via weights_sources + load_weights).
  • Upstream encode_instruction is exposed as encode_prompt (the vLLM-Omni convention, also hooked by the prompt-embed cache).
  • Text-to-image and single-reference TI2I inference share one native pipeline; CFG branches use vLLM-Omni's shared two-branch/N-branch parallel helpers.
  • Instruction rewriting, prompt tuning, and vision-token stripping are not ported.
  • BooguImagePipeline preserves the regular scheduler/CFG path, while BooguImageTurboPipeline selects the upstream few-step DMD student path.

SYSTEM_PROMPT_4_T2I_UNIFIED module-attribute

SYSTEM_PROMPT_4_T2I_UNIFIED = "You are a helpful assistant that generates high-quality images based on user instructions. The instructions are as follows."

SYSTEM_PROMPT_4_TI2I_UNIFIED module-attribute

SYSTEM_PROMPT_4_TI2I_UNIFIED = "Describe the key features of the input image (color, shape, size, texture, objects, background), then explain how the user's text instruction should alter or modify the image. Generate a new image that meets the user's requirements while maintaining consistency with the original input where appropriate."

logger module-attribute

logger = init_logger(__name__)

BooguImagePipeline

Bases: CFGParallelMixin, Module, ProgressBarMixin, SupportsComponentDiscovery, SupportImageInput

Boogu-Image text-to-image and image-editing (TI2I) pipeline.

Native vLLM-Omni implementation. A request with a reference image (edit / TI2I) is served by the same class as text-to-image; the reference latents and Qwen3VL image tokens are threaded through forward and the ported transformer's reference-image refiner path.

SYSTEM_PROMPT_4_I2I instance-attribute

SYSTEM_PROMPT_4_I2I = SYSTEM_PROMPT_4_TI2I_UNIFIED

SYSTEM_PROMPT_4_T2I instance-attribute

SYSTEM_PROMPT_4_T2I = SYSTEM_PROMPT_4_T2I_UNIFIED

SYSTEM_PROMPT_4_TI2I instance-attribute

SYSTEM_PROMPT_4_TI2I = SYSTEM_PROMPT_4_TI2I_UNIFIED

SYSTEM_PROMPT_DROP instance-attribute

SYSTEM_PROMPT_DROP = SYSTEM_PROMPT_4_TI2I_UNIFIED

color_format class-attribute

color_format: str = 'RGB'

default_sample_size instance-attribute

default_sample_size = 128

mllm instance-attribute

mllm = self._load_mllm(
    model,
    local_files_only,
    mllm_quant_config,
    boogu_subfolders,
)

od_config instance-attribute

od_config = od_config

processor instance-attribute

processor = Qwen3VLProcessor.from_pretrained(
    model,
    subfolder="processor",
    local_files_only=local_files_only,
    revision=od_config.revision,
)

scheduler instance-attribute

scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
    model,
    subfolder="scheduler",
    local_files_only=local_files_only,
    revision=od_config.revision,
)

support_image_input class-attribute

support_image_input: bool = True

supports_request_batch class-attribute instance-attribute

supports_request_batch = True

transformer instance-attribute

transformer = BooguImageTransformer2DModel(
    od_config=od_config,
    quant_config=transformer_quant_config,
    prefix="transformer",
)

vae instance-attribute

vae = from_pretrained_with_prefetch(
    AutoencoderKL.from_pretrained,
    model,
    subfolder="vae",
    prefetch_list=boogu_subfolders,
    local_files_only=local_files_only,
    revision=od_config.revision,
).to(self._execution_device)

vae_scale_factor instance-attribute

vae_scale_factor = (
    2 ** (len(self.vae.config.block_out_channels) - 1)
    if getattr(self, "vae", None)
    else 8
)

weights_sources instance-attribute

weights_sources = [
    DiffusersPipelineLoader.ComponentSource(
        model_or_path=od_config.model,
        subfolder="transformer",
        revision=od_config.revision,
        prefix="transformer.",
        fall_back_to_pt=True,
    )
]

combine_cfg_noise

combine_cfg_noise(
    positive_noise_pred: Tensor | tuple[Tensor, ...],
    negative_noise_pred: Tensor | tuple[Tensor, ...],
    true_cfg_scale: float,
    cfg_normalize: bool = False,
    kwargs: dict | None = None,
) -> Tensor

Preserve Boogu's sequential two-branch CFG operation order.

combine_multi_branch_cfg_noise

combine_multi_branch_cfg_noise(
    predictions: list[Tensor],
    true_cfg_scale: float | dict[str, float],
    cfg_normalize: bool = False,
) -> Tensor

Combine Boogu CFG branches using the original operation order.

Although the usual two- and three-branch CFG formulas can be rewritten algebraically, changing their floating-point operation order causes small per-step differences that accumulate across the denoise loop. Keep the sequential Boogu implementation's order so parallel branch combination does not introduce an additional source of numeric drift.

The two-branch formula is::

positive + (scale - 1) * (positive - negative)

The three-branch order is [positive_with_reference, negative_with_reference, negative_without_reference] and combines as::

positive_with_reference
+ (text_scale - 1) * (positive_with_reference - negative_with_reference)
+ (image_scale - 1) * (negative_with_reference - uncond)

encode_prompt

encode_prompt(
    prompt: str | list[str],
    do_classifier_free_guidance: bool = True,
    negative_prompt: str | list[str] | None = None,
    num_images_per_prompt: int = 1,
    device: device | None = None,
    prompt_embeds: Tensor | None = None,
    negative_prompt_embeds: Tensor | None = None,
    prompt_attention_mask: Tensor | None = None,
    negative_prompt_attention_mask: Tensor | None = None,
    max_sequence_length: int = 1280,
    truncate_instruction_sequence: bool = False,
    input_images: list[list[Image] | None] | None = None,
) -> tuple[Tensor, Tensor, Tensor | None, Tensor | None]

Encode prompt (and negative prompt for CFG) into Qwen3VL hidden states.

Port of upstream encode_instruction for text-to-image and the text-guided image-editing (TI2I) path. Reference images are attached to the positive instruction only (upstream default use_input_images_4_neg_instruct=False). Instruction rewriting, prompt tuning, and double-guidance empty instructions are not ported. The default max_sequence_length matches the upstream __call__ default (1280), not the upstream encode_instruction default (256).

Parameters:

Name Type Description Default
input_images list[list[Image] | None] | None

Per-sample list (outer length == batch size) of already-VLM-resized reference images, or None for pure text-to-image.

None

Returns:

Type Description
Tensor

``(prompt_embeds, prompt_attention_mask, negative_prompt_embeds,

Tensor

negative_prompt_attention_mask)`` where each embeds tensor has shape

Tensor | None

[batch_size * num_images_per_prompt, seq_len, dim]. The negative

Tensor | None

pair is None when do_classifier_free_guidance is off and no

tuple[Tensor, Tensor, Tensor | None, Tensor | None]

precomputed negative embeddings were passed.

forward

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

predict

predict(
    t,
    latents,
    instruction_embeds,
    freqs_real,
    instruction_attention_mask,
    ref_image_hidden_states=None,
)

One transformer velocity prediction (upstream predict).

ref_image_hidden_states is None for text-to-image, or the per-sample reference latents (list[list[Tensor[C, H, W]]]) for the image-editing path.

predict_noise

predict_noise(**kwargs) -> Tensor

Run one Boogu CFG branch through the native transformer.

prepare_latents

prepare_latents(
    batch_size,
    num_channels_latents,
    height,
    width,
    dtype,
    device,
    generator,
    latents=None,
)

Sample initial noise latents (upstream prepare_latents).

BooguImageTurboPipeline

Bases: BooguImagePipeline

Boogu-Image Turbo pipeline using the upstream few-step DMD semantics.

supports_request_batch class-attribute instance-attribute

supports_request_batch = False

get_boogu_image_post_process_func

get_boogu_image_post_process_func(
    od_config: OmniDiffusionConfig,
)

Build the post-process callable that converts decoded tensors to images.

Upstream BooguImageProcessor only customizes pre-processing; the postprocess path is inherited from the stock diffusers VaeImageProcessor, so we reuse it directly here.

get_boogu_image_pre_process_func

get_boogu_image_pre_process_func(
    od_config: OmniDiffusionConfig,
)

Build the pre-process callable for Boogu-Image reference (edit) input.

Text-to-image requests carry no image and are passed through unchanged (the Base checkpoint shares this pipeline class). Edit (TI2I) requests carry a single reference PIL image on prompt["multi_modal_data"]["image"]; it is resized twice — once for the Qwen3VL encoder (prompt_image) and once for the VAE reference latents (preprocessed_image) — and stashed in additional_information for forward to consume. Mirrors upstream preprocess_vlm_input_pil_images + prepare_image.

For a single reference image, upstream align_res (default True) derives the output resolution from the VAE-encoded reference dimensions, so the request height/width are overwritten accordingly.