vllm_omni.diffusion.models.sensenova_u1.pipeline_sensenova_u1 ¶
SenseNova-U1 Pipeline for vLLM-Omni.
SenseNova-U1 is a unified Qwen3-based model that uses Mixture-of-Tokenizers (MoT) attention for text-to-image generation via flow matching in patch space. It has no separate VAE or text encoder — the Qwen3 LLM itself serves as both the text encoder (via KV cache) and the denoising backbone (via MoT branches).
Key integration points: - Transformer layers ported with TP support (QKVParallelLinear, MergedColumnParallelLinear, RowParallelLinear) in sensenova_u1_transformer.py. - Vision model (NEOVisionModel) and FM modules kept as standard nn.Module since they are lightweight (no transformer blocks). - Weight loading uses stacked_params_mapping for fused QKV and gate_up.
SYSTEM_MESSAGE_FOR_GEN module-attribute ¶
SYSTEM_MESSAGE_FOR_GEN = "You are an image generation and editing assistant that accurately understands and executes user intent.\n\nYou support two modes:\n\n1. Think Mode:\nIf the task requires reasoning, you MUST start with a <think></think> block. Put all reasoning inside the block using plain text. DO NOT include any image tags. Keep it reasonable and directly useful for producing the final image.\n\n2. Non-Think Mode:\nIf no reasoning is needed, directly produce the final image.\n\nTask Types:\n\nA. Text-to-Image Generation:\n- Generate a high-quality image based on the user's description.\n- Ensure visual clarity, semantic consistency, and completeness.\n- DO NOT introduce elements that contradict or override the user's intent.\n\nB. Image Editing:\n- Use the provided image(s) as input or reference for modification or transformation.\n- The result can be an edited image or a new image based on the reference(s).\n- Preserve all unspecified attributes unless explicitly changed.\n\nGeneral Rules:\n- For any visible text in the image, follow the language specified for the rendered text in the user's description, not the language of the prompt. If no language is specified, use the user's input language."
NEOVisionEmbeddings ¶
Bases: Module
dense_embedding instance-attribute ¶
dense_embedding = nn.Conv2d(
self.embed_dim,
self.llm_embed_dim,
kernel_size=self.downsample_factor,
stride=self.downsample_factor,
)
patch_embedding instance-attribute ¶
patch_embedding = nn.Conv2d(
config.num_channels,
self.embed_dim,
kernel_size=self.patch_size,
stride=self.patch_size,
)
NEOVisionModel ¶
Bases: Module
SenseNovaU1DenoisingAdapter ¶
SenseNovaU1Pipeline ¶
Bases: Module, SupportsComponentDiscovery, DiffusionPipelineProfilerMixin, CFGParallelMixin, LoraLoaderMixin
SenseNova-U1 text-to-image and image-to-image pipeline for vllm-omni.
Builds the full model graph internally: - language_model: SenseNovaU1ForCausalLM (TP-aware) - vision_model: NEOVisionModel (understanding branch) - fm_modules: ModuleDict with vision_model_mot_gen, timestep_embedder, fm_head, etc.
img2img (image editing) is triggered when multi_modal_data["image"] is present in the prompt dict. The pipeline then uses triple KV caches (condition / img_condition / uncondition) with dual CFG (cfg_scale + img_cfg_scale).
denoising_transformer instance-attribute ¶
denoising_transformer = SenseNovaU1DenoisingAdapter(
self.language_model
)
fm_modules instance-attribute ¶
fm_modules = nn.ModuleDict(
{
"vision_model_mot_gen": vision_model_mot_gen,
"timestep_embedder": timestep_embedder,
"fm_head": fm_head,
}
)
img_context_token_id instance-attribute ¶
img_context_token_id = self.tokenizer.convert_tokens_to_ids(
IMG_CONTEXT_TOKEN
)
img_start_token_id instance-attribute ¶
img_start_token_id = self.tokenizer.convert_tokens_to_ids(
IMG_START_TOKEN
)
language_model instance-attribute ¶
language_model = SenseNovaU1ForCausalLM(
self.llm_cfg, prefix="language_model"
)
model_cfg instance-attribute ¶
model_cfg = SenseNovaU1Config.from_pretrained(
self.local_model_path
)
stacked_params_mapping class-attribute ¶
stacked_params_mapping: list[tuple[str, str, str | int]] = [
(".qkv_proj_mot_gen", ".q_proj_mot_gen", "q"),
(".qkv_proj_mot_gen", ".k_proj_mot_gen", "k"),
(".qkv_proj_mot_gen", ".v_proj_mot_gen", "v"),
(".qkv_proj", ".q_proj", "q"),
(".qkv_proj", ".k_proj", "k"),
(".qkv_proj", ".v_proj", "v"),
(".gate_up_proj", ".gate_proj", 0),
(".gate_up_proj", ".up_proj", 1),
]
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=self.local_model_path,
subfolder=None,
revision=od_config.revision,
prefix="",
fall_back_to_pt=False,
)
]
combine_cfg_noise ¶
combine_cfg_noise(
out_cond,
out_uncond,
cfg_scale,
cfg_norm,
kwargs: dict[str, Any] | None = None,
)
combine_multi_branch_cfg_noise ¶
load_lora_weights ¶
load_lora_weights(
pretrained_model_name_or_path: str | list[str],
adapter_name: str | None = None,
) -> None
Fuse a distilled few-step LoRA into the weights.
The checkpoints ship kohya lora_down/lora_up/alpha names, renamed here to the Diffusers names before load_lora_into_module routes each delta into its slice of the fused projections.
release_captured_graphs ¶
Drop the reused paged cache and the graphs captured against it.
Sleep level 2 discards the memory a capture recorded, so anything held across requests has to go with it. The next request rebuilds both.
TimestepEmbedder ¶
Bases: Module
mlp instance-attribute ¶
mlp = nn.Sequential(
nn.Linear(
frequency_embedding_size, hidden_size, bias=True
),
nn.SiLU(),
nn.Linear(hidden_size, hidden_size, bias=True),
)
get_sensenova_u1_post_process_func ¶
get_sensenova_u1_post_process_func(
od_config: OmniDiffusionConfig,
)