Skip to content

vllm_omni.diffusion.models.boogu_image

Modules:

Name Description
boogu_image_transformer
image_processor

Native port of the upstream BooguImageProcessor.

pipeline_boogu_image

Native vLLM-Omni pipeline for Boogu-Image-0.1.

scheduling_flow_match_euler_discrete_time_shifting

BooguImagePipeline

Bases: CFGParallelMixin, Module, ProgressBarMixin, SupportsComponentDiscovery, SupportImageInput

Boogu-Image text-to-image and image-editing (TI2I) pipeline.

Native vLLM-Omni implementation. A request with a reference image (edit / TI2I) is served by the same class as text-to-image; the reference latents and Qwen3VL image tokens are threaded through forward and the ported transformer's reference-image refiner path.

SYSTEM_PROMPT_4_I2I instance-attribute

SYSTEM_PROMPT_4_I2I = SYSTEM_PROMPT_4_TI2I_UNIFIED

SYSTEM_PROMPT_4_T2I instance-attribute

SYSTEM_PROMPT_4_T2I = SYSTEM_PROMPT_4_T2I_UNIFIED

SYSTEM_PROMPT_4_TI2I instance-attribute

SYSTEM_PROMPT_4_TI2I = SYSTEM_PROMPT_4_TI2I_UNIFIED

SYSTEM_PROMPT_DROP instance-attribute

SYSTEM_PROMPT_DROP = SYSTEM_PROMPT_4_TI2I_UNIFIED

color_format class-attribute

color_format: str = 'RGB'

default_sample_size instance-attribute

default_sample_size = 128

mllm instance-attribute

mllm = self._load_mllm(
    model,
    local_files_only,
    mllm_quant_config,
    boogu_subfolders,
)

od_config instance-attribute

od_config = od_config

processor instance-attribute

processor = Qwen3VLProcessor.from_pretrained(
    model,
    subfolder="processor",
    local_files_only=local_files_only,
    revision=od_config.revision,
)

scheduler instance-attribute

scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
    model,
    subfolder="scheduler",
    local_files_only=local_files_only,
    revision=od_config.revision,
)

support_image_input class-attribute

support_image_input: bool = True

supports_request_batch class-attribute instance-attribute

supports_request_batch = True

transformer instance-attribute

transformer = BooguImageTransformer2DModel(
    od_config=od_config,
    quant_config=transformer_quant_config,
    prefix="transformer",
)

vae instance-attribute

vae = from_pretrained_with_prefetch(
    AutoencoderKL.from_pretrained,
    model,
    subfolder="vae",
    prefetch_list=boogu_subfolders,
    local_files_only=local_files_only,
    revision=od_config.revision,
).to(self._execution_device)

vae_scale_factor instance-attribute

vae_scale_factor = (
    2 ** (len(self.vae.config.block_out_channels) - 1)
    if getattr(self, "vae", None)
    else 8
)

weights_sources instance-attribute

weights_sources = [
    DiffusersPipelineLoader.ComponentSource(
        model_or_path=od_config.model,
        subfolder="transformer",
        revision=od_config.revision,
        prefix="transformer.",
        fall_back_to_pt=True,
    )
]

combine_cfg_noise

combine_cfg_noise(
    positive_noise_pred: Tensor | tuple[Tensor, ...],
    negative_noise_pred: Tensor | tuple[Tensor, ...],
    true_cfg_scale: float,
    cfg_normalize: bool = False,
    kwargs: dict | None = None,
) -> Tensor

Preserve Boogu's sequential two-branch CFG operation order.

combine_multi_branch_cfg_noise

combine_multi_branch_cfg_noise(
    predictions: list[Tensor],
    true_cfg_scale: float | dict[str, float],
    cfg_normalize: bool = False,
) -> Tensor

Combine Boogu CFG branches using the original operation order.

Although the usual two- and three-branch CFG formulas can be rewritten algebraically, changing their floating-point operation order causes small per-step differences that accumulate across the denoise loop. Keep the sequential Boogu implementation's order so parallel branch combination does not introduce an additional source of numeric drift.

The two-branch formula is::

positive + (scale - 1) * (positive - negative)

The three-branch order is [positive_with_reference, negative_with_reference, negative_without_reference] and combines as::

positive_with_reference
+ (text_scale - 1) * (positive_with_reference - negative_with_reference)
+ (image_scale - 1) * (negative_with_reference - uncond)

encode_prompt

encode_prompt(
    prompt: str | list[str],
    do_classifier_free_guidance: bool = True,
    negative_prompt: str | list[str] | None = None,
    num_images_per_prompt: int = 1,
    device: device | None = None,
    prompt_embeds: Tensor | None = None,
    negative_prompt_embeds: Tensor | None = None,
    prompt_attention_mask: Tensor | None = None,
    negative_prompt_attention_mask: Tensor | None = None,
    max_sequence_length: int = 1280,
    truncate_instruction_sequence: bool = False,
    input_images: list[list[Image] | None] | None = None,
) -> tuple[Tensor, Tensor, Tensor | None, Tensor | None]

Encode prompt (and negative prompt for CFG) into Qwen3VL hidden states.

Port of upstream encode_instruction for text-to-image and the text-guided image-editing (TI2I) path. Reference images are attached to the positive instruction only (upstream default use_input_images_4_neg_instruct=False). Instruction rewriting, prompt tuning, and double-guidance empty instructions are not ported. The default max_sequence_length matches the upstream __call__ default (1280), not the upstream encode_instruction default (256).

Parameters:

Name Type Description Default
input_images list[list[Image] | None] | None

Per-sample list (outer length == batch size) of already-VLM-resized reference images, or None for pure text-to-image.

None

Returns:

Type Description
Tensor

``(prompt_embeds, prompt_attention_mask, negative_prompt_embeds,

Tensor

negative_prompt_attention_mask)`` where each embeds tensor has shape

Tensor | None

[batch_size * num_images_per_prompt, seq_len, dim]. The negative

Tensor | None

pair is None when do_classifier_free_guidance is off and no

tuple[Tensor, Tensor, Tensor | None, Tensor | None]

precomputed negative embeddings were passed.

forward

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

predict

predict(
    t,
    latents,
    instruction_embeds,
    freqs_real,
    instruction_attention_mask,
    ref_image_hidden_states=None,
)

One transformer velocity prediction (upstream predict).

ref_image_hidden_states is None for text-to-image, or the per-sample reference latents (list[list[Tensor[C, H, W]]]) for the image-editing path.

predict_noise

predict_noise(**kwargs) -> Tensor

Run one Boogu CFG branch through the native transformer.

prepare_latents

prepare_latents(
    batch_size,
    num_channels_latents,
    height,
    width,
    dtype,
    device,
    generator,
    latents=None,
)

Sample initial noise latents (upstream prepare_latents).

BooguImageProcessor

Bases: VaeImageProcessor

VaeImageProcessor variant with Boogu-Image pixel/side-length constraints.

Resizing never upscales (the ratio is clamped to <= 1) and always aligns the target height/width to multiples of vae_scale_factor.

max_pixels instance-attribute

max_pixels = max_pixels

max_side_length instance-attribute

max_side_length = max_side_length

get_new_height_width

get_new_height_width(
    image: Image | ndarray | Tensor,
    height: int | None = None,
    width: int | None = None,
    max_pixels: int | None = None,
    max_side_length: int | None = None,
) -> tuple[int, int]

Return target (height, width) after downscale + alignment.

Faithful port of upstream BooguImageProcessor.get_new_height_width.

preprocess

preprocess(
    image: PipelineImageInput,
    height: int | None = None,
    width: int | None = None,
    max_pixels: int | None = None,
    max_side_length: int | None = None,
    resize_mode: str = "default",
    crops_coords: tuple[int, int, int, int] | None = None,
) -> Tensor

Preprocess an image into a normalized [B, C, H, W] tensor.

Faithful port of upstream BooguImageProcessor.preprocess (PixArt-style downscale). Only the PIL branch is exercised by the native pipeline, but the numpy/tensor branches are kept for parity.

BooguImageTransformer2DModel

Bases: Module

Boogu-Image transformer with mixed stream topology.

Early layers use double-stream (dual-stream) processing, then switch to single-stream joint processing. Context/noise/reference-image refiner blocks run before the main stack.

axes_dim_rope instance-attribute

axes_dim_rope = axes_dim_rope

axes_lens instance-attribute

axes_lens = axes_lens

context_refiner instance-attribute

context_refiner = nn.ModuleList(
    [
        BooguImageContextRefinerTransformerBlock(
            hidden_size,
            num_attention_heads,
            num_kv_heads,
            multiple_of,
            ffn_dim_multiplier,
            norm_eps,
            modulation=False,
            quant_config=quant_config,
            prefix=_join_prefix(
                prefix, f"context_refiner.{i}"
            ),
        )
        for i in range(num_refiner_layers)
    ]
)

double_stream_layers instance-attribute

double_stream_layers = nn.ModuleList(
    [
        BooguImageDoubleStreamTransformerBlock(
            hidden_size,
            num_attention_heads,
            num_kv_heads,
            multiple_of,
            ffn_dim_multiplier,
            norm_eps,
            modulation=True,
            quant_config=quant_config,
            prefix=_join_prefix(
                prefix, f"double_stream_layers.{i}"
            ),
        )
        for i in range(num_double_stream_layers)
    ]
)

hidden_size instance-attribute

hidden_size = hidden_size

image_index_embedding instance-attribute

image_index_embedding = nn.Parameter(
    torch.randn(5, hidden_size)
)

in_channels instance-attribute

in_channels = in_channels

instruction_feature_configs instance-attribute

instruction_feature_configs = instruction_feature_configs

noise_refiner instance-attribute

noise_refiner = nn.ModuleList(
    [
        BooguImageNoiseRefinerTransformerBlock(
            hidden_size,
            num_attention_heads,
            num_kv_heads,
            multiple_of,
            ffn_dim_multiplier,
            norm_eps,
            modulation=True,
            quant_config=quant_config,
            prefix=_join_prefix(
                prefix, f"noise_refiner.{i}"
            ),
        )
        for i in range(num_refiner_layers)
    ]
)

norm_out instance-attribute

norm_out = LuminaLayerNormContinuous(
    embedding_dim=hidden_size,
    conditioning_embedding_dim=min(hidden_size, 1024),
    elementwise_affine=False,
    eps=1e-06,
    bias=True,
    out_dim=patch_size * patch_size * self.out_channels,
    quant_config=quant_config,
    prefix=_join_prefix(prefix, "norm_out"),
)

num_attention_heads instance-attribute

num_attention_heads = num_attention_heads

num_double_stream_layers instance-attribute

num_double_stream_layers = num_double_stream_layers

num_single_stream_layers instance-attribute

num_single_stream_layers = (
    num_layers - num_double_stream_layers
)

od_config instance-attribute

od_config = od_config

out_channels instance-attribute

out_channels = out_channels or in_channels

patch_size instance-attribute

patch_size = patch_size

preprocessed_instruction_feat_dim instance-attribute

preprocessed_instruction_feat_dim = (
    _cal_preprocessed_instruction_feat_dim(
        instruction_feature_configs
    )
)

ref_image_patch_embedder instance-attribute

ref_image_patch_embedder = nn.Linear(
    in_features=patch_size * patch_size * in_channels,
    out_features=hidden_size,
)

ref_image_refiner instance-attribute

ref_image_refiner = nn.ModuleList(
    [
        BooguImageRefImgRefinerTransformerBlock(
            hidden_size,
            num_attention_heads,
            num_kv_heads,
            multiple_of,
            ffn_dim_multiplier,
            norm_eps,
            modulation=True,
            quant_config=quant_config,
            prefix=_join_prefix(
                prefix, f"ref_image_refiner.{i}"
            ),
        )
        for i in range(num_refiner_layers)
    ]
)

rope_embedder instance-attribute

rope_embedder = BooguImageDoubleStreamRotaryPosEmbed(
    theta=10000,
    axes_dim=axes_dim_rope,
    axes_lens=axes_lens,
    patch_size=patch_size,
)

single_stream_layers instance-attribute

single_stream_layers = nn.ModuleList(
    [
        BooguImageSingleStreamTransformerBlock(
            hidden_size,
            num_attention_heads,
            num_kv_heads,
            multiple_of,
            ffn_dim_multiplier,
            norm_eps,
            modulation=True,
            quant_config=quant_config,
            prefix=_join_prefix(
                prefix, f"single_stream_layers.{i}"
            ),
        )
        for i in range(self.num_single_stream_layers)
    ]
)

stacked_params_mapping instance-attribute

stacked_params_mapping = list(_BOOGU_STACKED_PARAMS_MAPPING)

time_caption_embed instance-attribute

time_caption_embed = Lumina2CombinedTimestepCaptionEmbedding(
    hidden_size=hidden_size,
    instruction_feat_dim=self.preprocessed_instruction_feat_dim,
    norm_eps=norm_eps,
    timestep_scale=timestep_scale,
)

x_embedder instance-attribute

x_embedder = nn.Linear(
    in_features=patch_size * patch_size * in_channels,
    out_features=hidden_size,
)

flat_and_pad_to_seq

flat_and_pad_to_seq(hidden_states, ref_image_hidden_states)

Flatten patch tokens and pad to batched sequences.

Ported from upstream; for text-to-image ref_image_hidden_states is None and the reference-image branch collapses to zero-length.

forward

forward(
    hidden_states: Tensor | list[Tensor],
    timestep: Tensor,
    instruction_hidden_states: Tensor,
    freqs_real: RotaryFrequencyTables,
    instruction_attention_mask: Tensor,
    ref_image_hidden_states: list[list[Tensor]]
    | None = None,
) -> Tensor

Denoise one step: refiner -> double-stream -> fuse -> single-stream -> unpatchify.

Ported from upstream BooguImageTransformer2DModel.forward with the TeaCache/TaylorSeer/PEFT/gradient-checkpointing branches removed. Returns the velocity prediction as a [B, C, H_lat, W_lat] tensor.

img_patch_embed_and_refine

img_patch_embed_and_refine(
    hidden_states: Tensor,
    ref_image_hidden_states: Tensor,
    padded_img_mask: Tensor,
    padded_ref_img_mask: Tensor,
    noise_rotary_emb: RotaryEmbedding,
    ref_img_rotary_emb: RotaryEmbedding,
    l_effective_ref_img_len: list[list[int]],
    l_effective_img_len: list[int],
    temb: Tensor,
)

Embed image patches and run the refiner blocks.

The reference-image refiner is skipped when there are no reference-image tokens (text-to-image), which is numerically identical to upstream (the combined sequence only reads [:sum(ref_img_len)] = empty) while avoiding a degenerate zero-length attention.

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

Load diffusers-named checkpoint weights into the native module.

Name promotions relative to upstream (see step 8/10 findings):

  • *.img_instruct_attn.processor.{img,instruct}_{to_q,to_k,to_v} / {instruct,img}_out -> drop .processor (upstream keeps the joint-attention projections on the attention processor; the native module hosts them directly).
  • *.to_out.0.weight -> *.to_out.weight (diffusers wraps the output projection in a ModuleList; the native module uses a plain linear).

Packed Q/K/V and FFN gate/input matrices are folded onto the fused QKVParallelLinear / MergedColumnParallelLinear parameters using :attr:stacked_params_mapping — the same mapping loader consumers such as LoRA discovery and quantized weight loaders read.

A fused parameter is only reported as loaded once every one of its source matrices has arrived. The caller compares parameter names, so reporting e.g. to_qkv complete after the first shard would let a checkpoint carrying only to_q start up with the k/v slices left uninitialized.

preprocess_instruction_hidden_states

preprocess_instruction_hidden_states(
    raw_instruction_hidden_states,
)

Reduce the raw MLLM hidden states to the transformer feature dim.

Mirrors upstream preprocess_instruction_hidden_states: a single tensor passes through unchanged; a list of per-layer states is combined by concat or mean according to instruction_feature_configs.

BooguImageTurboPipeline

Bases: BooguImagePipeline

Boogu-Image Turbo pipeline using the upstream few-step DMD semantics.

supports_request_batch class-attribute instance-attribute

supports_request_batch = False

FlowMatchEulerDiscreteScheduler

Bases: SchedulerMixin, ConfigMixin

Euler scheduler with Boogu's training-consistent time shifting.

Timesteps run ascending 0 -> 1 (not descending sigmas as in stock diffusers), set_timesteps accepts a num_tokens argument for the dynamic-shift variants, and step() integrates with t_next - t directly.

This model inherits from [SchedulerMixin] and [ConfigMixin]. Check the superclass documentation for the generic methods the library implements for all schedulers such as loading and saving.

Parameters:

Name Type Description Default
num_train_timesteps `int`, defaults to 1000

The number of diffusion steps to train the model.

1000
do_shift `bool`, defaults to `True`

Whether to apply the training-consistent time shift in set_timesteps.

True
dynamic_time_shift `bool`, defaults to `True`

If True, the shift depends on the per-sample num_tokens; if False, it uses the configured seq_len.

True
time_shift_version `str`, defaults to `"v2"`

"v1" (logistic transform with a linear mu mapping) or "v2" (rational transform).

'v2'
seq_len `int`, *optional*

Token count used to compute the static shift when dynamic_time_shift=False (mirrors training).

None
base_shift `float`, defaults to 0.5

v1 linear mapping lower bound (matches training defaults).

0.5
max_shift `float`, defaults to 1.15

v1 linear mapping upper bound (matches training defaults).

1.15

begin_index property

begin_index

The index for the first timestep. It should be set from pipeline with set_begin_index method.

order class-attribute instance-attribute

order = 1

step_index property

step_index

The index counter for current timestep. It will increase 1 after each scheduler step.

time_shift_v2_scaling_factor instance-attribute

time_shift_v2_scaling_factor = (
    time_shift_v2_half_scaling_factor * 2
)

timesteps instance-attribute

timesteps = timesteps

index_for_timestep

index_for_timestep(timestep, schedule_timesteps=None)

set_begin_index

set_begin_index(begin_index: int = 0)

Sets the begin index for the scheduler. This function should be run from pipeline before the inference.

Parameters:

Name Type Description Default
begin_index `int`

The begin index for the scheduler.

0

set_timesteps

set_timesteps(
    num_inference_steps: int | None = None,
    device: str | device | None = None,
    timesteps: list[float] | None = None,
    num_tokens: int | None = None,
)

Sets the discrete timesteps used for the diffusion chain (to be run before inference).

Parameters:

Name Type Description Default
num_inference_steps `int`

The number of diffusion steps used when generating samples with a pre-trained model.

None
device `str` or `torch.device`, *optional*

The device to which the timesteps should be moved to. If None, the timesteps are not moved.

None
timesteps `list[float]`, *optional*

Custom timesteps to use. If provided, num_inference_steps is ignored.

None
num_tokens `int`, *optional*

Per-sample token count, used by the dynamic time-shift variants.

None

step

step(
    model_output: FloatTensor,
    timestep: float | FloatTensor,
    sample: FloatTensor,
    generator: Generator | None = None,
    return_dict: bool = True,
) -> FlowMatchEulerDiscreteSchedulerOutput | tuple

Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion process from the learned model outputs (most often the predicted noise).

Parameters:

Name Type Description Default
model_output `torch.FloatTensor`

The direct output from learned diffusion model.

required
timestep `float`

The current discrete timestep in the diffusion chain.

required
sample `torch.FloatTensor`

A current instance of a sample created by the diffusion process.

required
generator `torch.Generator`, *optional*

A random number generator.

None
return_dict `bool`

Whether or not to return a [FlowMatchEulerDiscreteSchedulerOutput] or tuple.

True

Returns:

Type Description
FlowMatchEulerDiscreteSchedulerOutput | tuple

[FlowMatchEulerDiscreteSchedulerOutput] or tuple: If return_dict is True, [FlowMatchEulerDiscreteSchedulerOutput] is returned, otherwise a tuple is returned where the first element is the sample tensor.

get_boogu_image_pre_process_func

get_boogu_image_pre_process_func(
    od_config: OmniDiffusionConfig,
)

Build the pre-process callable for Boogu-Image reference (edit) input.

Text-to-image requests carry no image and are passed through unchanged (the Base checkpoint shares this pipeline class). Edit (TI2I) requests carry a single reference PIL image on prompt["multi_modal_data"]["image"]; it is resized twice — once for the Qwen3VL encoder (prompt_image) and once for the VAE reference latents (preprocessed_image) — and stashed in additional_information for forward to consume. Mirrors upstream preprocess_vlm_input_pil_images + prepare_image.

For a single reference image, upstream align_res (default True) derives the output resolution from the VAE-encoded reference dimensions, so the request height/width are overwritten accordingly.