vllm_omni.diffusion.models.boogu_image ¶
Modules:
| Name | Description |
|---|---|
boogu_image_transformer | |
image_processor | Native port of the upstream |
pipeline_boogu_image | Native vLLM-Omni pipeline for Boogu-Image-0.1. |
scheduling_flow_match_euler_discrete_time_shifting | |
BooguImagePipeline ¶
Bases: CFGParallelMixin, Module, ProgressBarMixin, SupportsComponentDiscovery, SupportImageInput
Boogu-Image text-to-image and image-editing (TI2I) pipeline.
Native vLLM-Omni implementation. A request with a reference image (edit / TI2I) is served by the same class as text-to-image; the reference latents and Qwen3VL image tokens are threaded through forward and the ported transformer's reference-image refiner path.
mllm instance-attribute ¶
processor instance-attribute ¶
processor = Qwen3VLProcessor.from_pretrained(
model,
subfolder="processor",
local_files_only=local_files_only,
revision=od_config.revision,
)
scheduler instance-attribute ¶
scheduler = FlowMatchEulerDiscreteScheduler.from_pretrained(
model,
subfolder="scheduler",
local_files_only=local_files_only,
revision=od_config.revision,
)
transformer instance-attribute ¶
transformer = BooguImageTransformer2DModel(
od_config=od_config,
quant_config=transformer_quant_config,
prefix="transformer",
)
vae instance-attribute ¶
vae = from_pretrained_with_prefetch(
AutoencoderKL.from_pretrained,
model,
subfolder="vae",
prefetch_list=boogu_subfolders,
local_files_only=local_files_only,
revision=od_config.revision,
).to(self._execution_device)
vae_scale_factor instance-attribute ¶
vae_scale_factor = (
2 ** (len(self.vae.config.block_out_channels) - 1)
if getattr(self, "vae", None)
else 8
)
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=od_config.model,
subfolder="transformer",
revision=od_config.revision,
prefix="transformer.",
fall_back_to_pt=True,
)
]
combine_cfg_noise ¶
combine_cfg_noise(
positive_noise_pred: Tensor | tuple[Tensor, ...],
negative_noise_pred: Tensor | tuple[Tensor, ...],
true_cfg_scale: float,
cfg_normalize: bool = False,
kwargs: dict | None = None,
) -> Tensor
Preserve Boogu's sequential two-branch CFG operation order.
combine_multi_branch_cfg_noise ¶
combine_multi_branch_cfg_noise(
predictions: list[Tensor],
true_cfg_scale: float | dict[str, float],
cfg_normalize: bool = False,
) -> Tensor
Combine Boogu CFG branches using the original operation order.
Although the usual two- and three-branch CFG formulas can be rewritten algebraically, changing their floating-point operation order causes small per-step differences that accumulate across the denoise loop. Keep the sequential Boogu implementation's order so parallel branch combination does not introduce an additional source of numeric drift.
The two-branch formula is::
positive + (scale - 1) * (positive - negative)
The three-branch order is [positive_with_reference, negative_with_reference, negative_without_reference] and combines as::
positive_with_reference
+ (text_scale - 1) * (positive_with_reference - negative_with_reference)
+ (image_scale - 1) * (negative_with_reference - uncond)
encode_prompt ¶
encode_prompt(
prompt: str | list[str],
do_classifier_free_guidance: bool = True,
negative_prompt: str | list[str] | None = None,
num_images_per_prompt: int = 1,
device: device | None = None,
prompt_embeds: Tensor | None = None,
negative_prompt_embeds: Tensor | None = None,
prompt_attention_mask: Tensor | None = None,
negative_prompt_attention_mask: Tensor | None = None,
max_sequence_length: int = 1280,
truncate_instruction_sequence: bool = False,
input_images: list[list[Image] | None] | None = None,
) -> tuple[Tensor, Tensor, Tensor | None, Tensor | None]
Encode prompt (and negative prompt for CFG) into Qwen3VL hidden states.
Port of upstream encode_instruction for text-to-image and the text-guided image-editing (TI2I) path. Reference images are attached to the positive instruction only (upstream default use_input_images_4_neg_instruct=False). Instruction rewriting, prompt tuning, and double-guidance empty instructions are not ported. The default max_sequence_length matches the upstream __call__ default (1280), not the upstream encode_instruction default (256).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input_images | list[list[Image] | None] | None | Per-sample list (outer length == batch size) of already-VLM-resized reference images, or | None |
Returns:
| Type | Description |
|---|---|
Tensor | ``(prompt_embeds, prompt_attention_mask, negative_prompt_embeds, |
Tensor | negative_prompt_attention_mask)`` where each embeds tensor has shape |
Tensor | None |
|
Tensor | None | pair is |
tuple[Tensor, Tensor, Tensor | None, Tensor | None] | precomputed negative embeddings were passed. |
predict ¶
predict(
t,
latents,
instruction_embeds,
freqs_real,
instruction_attention_mask,
ref_image_hidden_states=None,
)
One transformer velocity prediction (upstream predict).
ref_image_hidden_states is None for text-to-image, or the per-sample reference latents (list[list[Tensor[C, H, W]]]) for the image-editing path.
predict_noise ¶
Run one Boogu CFG branch through the native transformer.
prepare_latents ¶
prepare_latents(
batch_size,
num_channels_latents,
height,
width,
dtype,
device,
generator,
latents=None,
)
Sample initial noise latents (upstream prepare_latents).
BooguImageProcessor ¶
Bases: VaeImageProcessor
VaeImageProcessor variant with Boogu-Image pixel/side-length constraints.
Resizing never upscales (the ratio is clamped to <= 1) and always aligns the target height/width to multiples of vae_scale_factor.
get_new_height_width ¶
get_new_height_width(
image: Image | ndarray | Tensor,
height: int | None = None,
width: int | None = None,
max_pixels: int | None = None,
max_side_length: int | None = None,
) -> tuple[int, int]
Return target (height, width) after downscale + alignment.
Faithful port of upstream BooguImageProcessor.get_new_height_width.
preprocess ¶
preprocess(
image: PipelineImageInput,
height: int | None = None,
width: int | None = None,
max_pixels: int | None = None,
max_side_length: int | None = None,
resize_mode: str = "default",
crops_coords: tuple[int, int, int, int] | None = None,
) -> Tensor
Preprocess an image into a normalized [B, C, H, W] tensor.
Faithful port of upstream BooguImageProcessor.preprocess (PixArt-style downscale). Only the PIL branch is exercised by the native pipeline, but the numpy/tensor branches are kept for parity.
BooguImageTransformer2DModel ¶
Bases: Module
Boogu-Image transformer with mixed stream topology.
Early layers use double-stream (dual-stream) processing, then switch to single-stream joint processing. Context/noise/reference-image refiner blocks run before the main stack.
context_refiner instance-attribute ¶
context_refiner = nn.ModuleList(
[
BooguImageContextRefinerTransformerBlock(
hidden_size,
num_attention_heads,
num_kv_heads,
multiple_of,
ffn_dim_multiplier,
norm_eps,
modulation=False,
quant_config=quant_config,
prefix=_join_prefix(
prefix, f"context_refiner.{i}"
),
)
for i in range(num_refiner_layers)
]
)
double_stream_layers instance-attribute ¶
double_stream_layers = nn.ModuleList(
[
BooguImageDoubleStreamTransformerBlock(
hidden_size,
num_attention_heads,
num_kv_heads,
multiple_of,
ffn_dim_multiplier,
norm_eps,
modulation=True,
quant_config=quant_config,
prefix=_join_prefix(
prefix, f"double_stream_layers.{i}"
),
)
for i in range(num_double_stream_layers)
]
)
image_index_embedding instance-attribute ¶
instruction_feature_configs instance-attribute ¶
noise_refiner instance-attribute ¶
noise_refiner = nn.ModuleList(
[
BooguImageNoiseRefinerTransformerBlock(
hidden_size,
num_attention_heads,
num_kv_heads,
multiple_of,
ffn_dim_multiplier,
norm_eps,
modulation=True,
quant_config=quant_config,
prefix=_join_prefix(
prefix, f"noise_refiner.{i}"
),
)
for i in range(num_refiner_layers)
]
)
norm_out instance-attribute ¶
norm_out = LuminaLayerNormContinuous(
embedding_dim=hidden_size,
conditioning_embedding_dim=min(hidden_size, 1024),
elementwise_affine=False,
eps=1e-06,
bias=True,
out_dim=patch_size * patch_size * self.out_channels,
quant_config=quant_config,
prefix=_join_prefix(prefix, "norm_out"),
)
num_single_stream_layers instance-attribute ¶
preprocessed_instruction_feat_dim instance-attribute ¶
preprocessed_instruction_feat_dim = (
_cal_preprocessed_instruction_feat_dim(
instruction_feature_configs
)
)
ref_image_patch_embedder instance-attribute ¶
ref_image_patch_embedder = nn.Linear(
in_features=patch_size * patch_size * in_channels,
out_features=hidden_size,
)
ref_image_refiner instance-attribute ¶
ref_image_refiner = nn.ModuleList(
[
BooguImageRefImgRefinerTransformerBlock(
hidden_size,
num_attention_heads,
num_kv_heads,
multiple_of,
ffn_dim_multiplier,
norm_eps,
modulation=True,
quant_config=quant_config,
prefix=_join_prefix(
prefix, f"ref_image_refiner.{i}"
),
)
for i in range(num_refiner_layers)
]
)
rope_embedder instance-attribute ¶
rope_embedder = BooguImageDoubleStreamRotaryPosEmbed(
theta=10000,
axes_dim=axes_dim_rope,
axes_lens=axes_lens,
patch_size=patch_size,
)
single_stream_layers instance-attribute ¶
single_stream_layers = nn.ModuleList(
[
BooguImageSingleStreamTransformerBlock(
hidden_size,
num_attention_heads,
num_kv_heads,
multiple_of,
ffn_dim_multiplier,
norm_eps,
modulation=True,
quant_config=quant_config,
prefix=_join_prefix(
prefix, f"single_stream_layers.{i}"
),
)
for i in range(self.num_single_stream_layers)
]
)
stacked_params_mapping instance-attribute ¶
stacked_params_mapping = list(_BOOGU_STACKED_PARAMS_MAPPING)
time_caption_embed instance-attribute ¶
time_caption_embed = Lumina2CombinedTimestepCaptionEmbedding(
hidden_size=hidden_size,
instruction_feat_dim=self.preprocessed_instruction_feat_dim,
norm_eps=norm_eps,
timestep_scale=timestep_scale,
)
x_embedder instance-attribute ¶
x_embedder = nn.Linear(
in_features=patch_size * patch_size * in_channels,
out_features=hidden_size,
)
flat_and_pad_to_seq ¶
Flatten patch tokens and pad to batched sequences.
Ported from upstream; for text-to-image ref_image_hidden_states is None and the reference-image branch collapses to zero-length.
forward ¶
forward(
hidden_states: Tensor | list[Tensor],
timestep: Tensor,
instruction_hidden_states: Tensor,
freqs_real: RotaryFrequencyTables,
instruction_attention_mask: Tensor,
ref_image_hidden_states: list[list[Tensor]]
| None = None,
) -> Tensor
Denoise one step: refiner -> double-stream -> fuse -> single-stream -> unpatchify.
Ported from upstream BooguImageTransformer2DModel.forward with the TeaCache/TaylorSeer/PEFT/gradient-checkpointing branches removed. Returns the velocity prediction as a [B, C, H_lat, W_lat] tensor.
img_patch_embed_and_refine ¶
img_patch_embed_and_refine(
hidden_states: Tensor,
ref_image_hidden_states: Tensor,
padded_img_mask: Tensor,
padded_ref_img_mask: Tensor,
noise_rotary_emb: RotaryEmbedding,
ref_img_rotary_emb: RotaryEmbedding,
l_effective_ref_img_len: list[list[int]],
l_effective_img_len: list[int],
temb: Tensor,
)
Embed image patches and run the refiner blocks.
The reference-image refiner is skipped when there are no reference-image tokens (text-to-image), which is numerically identical to upstream (the combined sequence only reads [:sum(ref_img_len)] = empty) while avoiding a degenerate zero-length attention.
load_weights ¶
Load diffusers-named checkpoint weights into the native module.
Name promotions relative to upstream (see step 8/10 findings):
*.img_instruct_attn.processor.{img,instruct}_{to_q,to_k,to_v}/{instruct,img}_out-> drop.processor(upstream keeps the joint-attention projections on the attention processor; the native module hosts them directly).*.to_out.0.weight->*.to_out.weight(diffusers wraps the output projection in aModuleList; the native module uses a plain linear).
Packed Q/K/V and FFN gate/input matrices are folded onto the fused QKVParallelLinear / MergedColumnParallelLinear parameters using :attr:stacked_params_mapping — the same mapping loader consumers such as LoRA discovery and quantized weight loaders read.
A fused parameter is only reported as loaded once every one of its source matrices has arrived. The caller compares parameter names, so reporting e.g. to_qkv complete after the first shard would let a checkpoint carrying only to_q start up with the k/v slices left uninitialized.
preprocess_instruction_hidden_states ¶
Reduce the raw MLLM hidden states to the transformer feature dim.
Mirrors upstream preprocess_instruction_hidden_states: a single tensor passes through unchanged; a list of per-layer states is combined by concat or mean according to instruction_feature_configs.
BooguImageTurboPipeline ¶
FlowMatchEulerDiscreteScheduler ¶
Bases: SchedulerMixin, ConfigMixin
Euler scheduler with Boogu's training-consistent time shifting.
Timesteps run ascending 0 -> 1 (not descending sigmas as in stock diffusers), set_timesteps accepts a num_tokens argument for the dynamic-shift variants, and step() integrates with t_next - t directly.
This model inherits from [SchedulerMixin] and [ConfigMixin]. Check the superclass documentation for the generic methods the library implements for all schedulers such as loading and saving.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_train_timesteps | `int`, defaults to 1000 | The number of diffusion steps to train the model. | 1000 |
do_shift | `bool`, defaults to `True` | Whether to apply the training-consistent time shift in | True |
dynamic_time_shift | `bool`, defaults to `True` | If | True |
time_shift_version | `str`, defaults to `"v2"` |
| 'v2' |
seq_len | `int`, *optional* | Token count used to compute the static shift when | None |
base_shift | `float`, defaults to 0.5 | v1 linear mapping lower bound (matches training defaults). | 0.5 |
max_shift | `float`, defaults to 1.15 | v1 linear mapping upper bound (matches training defaults). | 1.15 |
begin_index property ¶
The index for the first timestep. It should be set from pipeline with set_begin_index method.
step_index property ¶
The index counter for current timestep. It will increase 1 after each scheduler step.
time_shift_v2_scaling_factor instance-attribute ¶
set_begin_index ¶
set_begin_index(begin_index: int = 0)
Sets the begin index for the scheduler. This function should be run from pipeline before the inference.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
begin_index | `int` | The begin index for the scheduler. | 0 |
set_timesteps ¶
set_timesteps(
num_inference_steps: int | None = None,
device: str | device | None = None,
timesteps: list[float] | None = None,
num_tokens: int | None = None,
)
Sets the discrete timesteps used for the diffusion chain (to be run before inference).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
num_inference_steps | `int` | The number of diffusion steps used when generating samples with a pre-trained model. | None |
device | `str` or `torch.device`, *optional* | The device to which the timesteps should be moved to. If | None |
timesteps | `list[float]`, *optional* | Custom timesteps to use. If provided, | None |
num_tokens | `int`, *optional* | Per-sample token count, used by the dynamic time-shift variants. | None |
step ¶
step(
model_output: FloatTensor,
timestep: float | FloatTensor,
sample: FloatTensor,
generator: Generator | None = None,
return_dict: bool = True,
) -> FlowMatchEulerDiscreteSchedulerOutput | tuple
Predict the sample from the previous timestep by reversing the SDE. This function propagates the diffusion process from the learned model outputs (most often the predicted noise).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
model_output | `torch.FloatTensor` | The direct output from learned diffusion model. | required |
timestep | `float` | The current discrete timestep in the diffusion chain. | required |
sample | `torch.FloatTensor` | A current instance of a sample created by the diffusion process. | required |
generator | `torch.Generator`, *optional* | A random number generator. | None |
return_dict | `bool` | Whether or not to return a [ | True |
Returns:
| Type | Description |
|---|---|
FlowMatchEulerDiscreteSchedulerOutput | tuple | [ |
get_boogu_image_pre_process_func ¶
get_boogu_image_pre_process_func(
od_config: OmniDiffusionConfig,
)
Build the pre-process callable for Boogu-Image reference (edit) input.
Text-to-image requests carry no image and are passed through unchanged (the Base checkpoint shares this pipeline class). Edit (TI2I) requests carry a single reference PIL image on prompt["multi_modal_data"]["image"]; it is resized twice — once for the Qwen3VL encoder (prompt_image) and once for the VAE reference latents (preprocessed_image) — and stashed in additional_information for forward to consume. Mirrors upstream preprocess_vlm_input_pil_images + prepare_image.
For a single reference image, upstream align_res (default True) derives the output resolution from the VAE-encoded reference dimensions, so the request height/width are overwritten accordingly.