vllm_omni.diffusion.models.sensenova_u1 ¶
Modules:
| Name | Description |
|---|---|
fused_rmsnorm_rope | |
paged_decode | Paged KV cache for SenseNova-U1 autoregressive decode, and a CUDA graph over it. |
pipeline_sensenova_u1 | SenseNova-U1 Pipeline for vLLM-Omni. |
sensenova_u1_transformer | Qwen3 LLM with Mixture-of-Tokenizers (MoT) for SenseNova-U1. |
SenseNovaU1Pipeline ¶
Bases: Module, SupportsComponentDiscovery, DiffusionPipelineProfilerMixin, CFGParallelMixin, LoraLoaderMixin
SenseNova-U1 text-to-image and image-to-image pipeline for vllm-omni.
Builds the full model graph internally: - language_model: SenseNovaU1ForCausalLM (TP-aware) - vision_model: NEOVisionModel (understanding branch) - fm_modules: ModuleDict with vision_model_mot_gen, timestep_embedder, fm_head, etc.
img2img (image editing) is triggered when multi_modal_data["image"] is present in the prompt dict. The pipeline then uses triple KV caches (condition / img_condition / uncondition) with dual CFG (cfg_scale + img_cfg_scale).
denoising_transformer instance-attribute ¶
denoising_transformer = SenseNovaU1DenoisingAdapter(
self.language_model
)
fm_modules instance-attribute ¶
fm_modules = nn.ModuleDict(
{
"vision_model_mot_gen": vision_model_mot_gen,
"timestep_embedder": timestep_embedder,
"fm_head": fm_head,
}
)
img_context_token_id instance-attribute ¶
img_context_token_id = self.tokenizer.convert_tokens_to_ids(
IMG_CONTEXT_TOKEN
)
img_start_token_id instance-attribute ¶
img_start_token_id = self.tokenizer.convert_tokens_to_ids(
IMG_START_TOKEN
)
language_model instance-attribute ¶
language_model = SenseNovaU1ForCausalLM(
self.llm_cfg, prefix="language_model"
)
model_cfg instance-attribute ¶
model_cfg = SenseNovaU1Config.from_pretrained(
self.local_model_path
)
stacked_params_mapping class-attribute ¶
stacked_params_mapping: list[tuple[str, str, str | int]] = [
(".qkv_proj_mot_gen", ".q_proj_mot_gen", "q"),
(".qkv_proj_mot_gen", ".k_proj_mot_gen", "k"),
(".qkv_proj_mot_gen", ".v_proj_mot_gen", "v"),
(".qkv_proj", ".q_proj", "q"),
(".qkv_proj", ".k_proj", "k"),
(".qkv_proj", ".v_proj", "v"),
(".gate_up_proj", ".gate_proj", 0),
(".gate_up_proj", ".up_proj", 1),
]
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=self.local_model_path,
subfolder=None,
revision=od_config.revision,
prefix="",
fall_back_to_pt=False,
)
]
combine_cfg_noise ¶
combine_cfg_noise(
out_cond,
out_uncond,
cfg_scale,
cfg_norm,
kwargs: dict[str, Any] | None = None,
)
combine_multi_branch_cfg_noise ¶
load_lora_weights ¶
load_lora_weights(
pretrained_model_name_or_path: str | list[str],
adapter_name: str | None = None,
) -> None
Fuse a distilled few-step LoRA into the weights.
The checkpoints ship kohya lora_down/lora_up/alpha names, renamed here to the Diffusers names before load_lora_into_module routes each delta into its slice of the fused projections.
release_captured_graphs ¶
Drop the reused paged cache and the graphs captured against it.
Sleep level 2 discards the memory a capture recorded, so anything held across requests has to go with it. The next request rebuilds both.
get_sensenova_u1_post_process_func ¶
get_sensenova_u1_post_process_func(
od_config: OmniDiffusionConfig,
)