vllm_omni.diffusion.models.boogu_image.pipeline_boogu_image ¶
Native vLLM-Omni pipeline for Boogu-Image-0.1.
Ported from the upstream boogu package (boogu/pipelines/boogu/pipeline_boogu.py) with the following changes:
- Diffusers
DiffusionPipeline/register_modulesmachinery replaced by a plainnn.Moduleconstructed fromOmniDiffusionConfig(components are loaded from the checkpoint subfolders; transformer weights arrive later viaweights_sources+load_weights). - Upstream
encode_instructionis exposed asencode_prompt(the vLLM-Omni convention, also hooked by the prompt-embed cache). - Text-to-image and single-reference TI2I inference share one native pipeline; CFG branches use vLLM-Omni's shared two-branch/N-branch parallel helpers.
- Instruction rewriting, prompt tuning, and vision-token stripping are not ported.
BooguImagePipelinepreserves the regular scheduler/CFG path, whileBooguImageTurboPipelineselects the upstream few-step DMD student path.
SYSTEM_PROMPT_4_T2I_UNIFIED module-attribute ¶
SYSTEM_PROMPT_4_T2I_UNIFIED = "You are a helpful assistant that generates high-quality images based on user instructions. The instructions are as follows."
SYSTEM_PROMPT_4_TI2I_UNIFIED module-attribute ¶
SYSTEM_PROMPT_4_TI2I_UNIFIED = "Describe the key features of the input image (color, shape, size, texture, objects, background), then explain how the user's text instruction should alter or modify the image. Generate a new image that meets the user's requirements while maintaining consistency with the original input where appropriate."
BooguImagePipeline ¶
Bases: CFGParallelMixin, Module, ProgressBarMixin, SupportsComponentDiscovery, SupportImageInput
Boogu-Image text-to-image and image-editing (TI2I) pipeline.
Native vLLM-Omni implementation. A request with a reference image (edit / TI2I) is served by the same class as text-to-image; the reference latents and Qwen3VL image tokens are threaded through forward and the ported transformer's reference-image refiner path.
mllm instance-attribute ¶
processor instance-attribute ¶
processor = Qwen3VLProcessor.from_pretrained(
model,
subfolder="processor",
local_files_only=local_files_only,
revision=od_config.revision,
)
scheduler instance-attribute ¶
scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
model,
subfolder="scheduler",
local_files_only=local_files_only,
revision=od_config.revision,
)
transformer instance-attribute ¶
transformer = BooguImageTransformer2DModel(
od_config=od_config,
quant_config=transformer_quant_config,
prefix="transformer",
)
vae instance-attribute ¶
vae = from_pretrained_with_prefetch(
AutoencoderKL.from_pretrained,
model,
subfolder="vae",
prefetch_list=boogu_subfolders,
local_files_only=local_files_only,
revision=od_config.revision,
).to(self._execution_device)
vae_scale_factor instance-attribute ¶
vae_scale_factor = (
2 ** (len(self.vae.config.block_out_channels) - 1)
if getattr(self, "vae", None)
else 8
)
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=od_config.model,
subfolder="transformer",
revision=od_config.revision,
prefix="transformer.",
fall_back_to_pt=True,
)
]
combine_cfg_noise ¶
combine_cfg_noise(
positive_noise_pred: Tensor | tuple[Tensor, ...],
negative_noise_pred: Tensor | tuple[Tensor, ...],
true_cfg_scale: float,
cfg_normalize: bool = False,
kwargs: dict | None = None,
) -> Tensor
Preserve Boogu's sequential two-branch CFG operation order.
combine_multi_branch_cfg_noise ¶
combine_multi_branch_cfg_noise(
predictions: list[Tensor],
true_cfg_scale: float | dict[str, float],
cfg_normalize: bool = False,
) -> Tensor
Combine Boogu CFG branches using the original operation order.
Although the usual two- and three-branch CFG formulas can be rewritten algebraically, changing their floating-point operation order causes small per-step differences that accumulate across the denoise loop. Keep the sequential Boogu implementation's order so parallel branch combination does not introduce an additional source of numeric drift.
The two-branch formula is::
positive + (scale - 1) * (positive - negative)
The three-branch order is [positive_with_reference, negative_with_reference, negative_without_reference] and combines as::
positive_with_reference
+ (text_scale - 1) * (positive_with_reference - negative_with_reference)
+ (image_scale - 1) * (negative_with_reference - uncond)
encode_prompt ¶
encode_prompt(
prompt: str | list[str],
do_classifier_free_guidance: bool = True,
negative_prompt: str | list[str] | None = None,
num_images_per_prompt: int = 1,
device: device | None = None,
prompt_embeds: Tensor | None = None,
negative_prompt_embeds: Tensor | None = None,
prompt_attention_mask: Tensor | None = None,
negative_prompt_attention_mask: Tensor | None = None,
max_sequence_length: int = 1280,
truncate_instruction_sequence: bool = False,
input_images: list[list[Image] | None] | None = None,
) -> tuple[Tensor, Tensor, Tensor | None, Tensor | None]
Encode prompt (and negative prompt for CFG) into Qwen3VL hidden states.
Port of upstream encode_instruction for text-to-image and the text-guided image-editing (TI2I) path. Reference images are attached to the positive instruction only (upstream default use_input_images_4_neg_instruct=False). Instruction rewriting, prompt tuning, and double-guidance empty instructions are not ported. The default max_sequence_length matches the upstream __call__ default (1280), not the upstream encode_instruction default (256).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_images | list[list[Image] | None] | None | Per-sample list (outer length == batch size) of already-VLM-resized reference images, or | None |
Returns:
| Type | Description |
|---|---|
Tensor | ``(prompt_embeds, prompt_attention_mask, negative_prompt_embeds, |
Tensor | negative_prompt_attention_mask)`` where each embeds tensor has shape |
Tensor | None |
|
Tensor | None | pair is |
tuple[Tensor, Tensor, Tensor | None, Tensor | None] | precomputed negative embeddings were passed. |
predict ¶
predict(
t,
latents,
instruction_embeds,
freqs_real,
instruction_attention_mask,
ref_image_hidden_states=None,
)
One transformer velocity prediction (upstream predict).
ref_image_hidden_states is None for text-to-image, or the per-sample reference latents (list[list[Tensor[C, H, W]]]) for the image-editing path.
predict_noise ¶
Run one Boogu CFG branch through the native transformer.
prepare_latents ¶
prepare_latents(
batch_size,
num_channels_latents,
height,
width,
dtype,
device,
generator,
latents=None,
)
Sample initial noise latents (upstream prepare_latents).
BooguImageTurboPipeline ¶
get_boogu_image_post_process_func ¶
get_boogu_image_post_process_func(
od_config: OmniDiffusionConfig,
)
Build the post-process callable that converts decoded tensors to images.
Upstream BooguImageProcessor only customizes pre-processing; the postprocess path is inherited from the stock diffusers VaeImageProcessor, so we reuse it directly here.
get_boogu_image_pre_process_func ¶
get_boogu_image_pre_process_func(
od_config: OmniDiffusionConfig,
)
Build the pre-process callable for Boogu-Image reference (edit) input.
Text-to-image requests carry no image and are passed through unchanged (the Base checkpoint shares this pipeline class). Edit (TI2I) requests carry a single reference PIL image on prompt["multi_modal_data"]["image"]; it is resized twice — once for the Qwen3VL encoder (prompt_image) and once for the VAE reference latents (preprocessed_image) — and stashed in additional_information for forward to consume. Mirrors upstream preprocess_vlm_input_pil_images + prepare_image.
For a single reference image, upstream align_res (default True) derives the output resolution from the VAE-encoded reference dimensions, so the request height/width are overwritten accordingly.