Skip to content

vllm_omni.diffusion.models.sensenova_u1

Modules:

Name Description
fused_rmsnorm_rope
paged_decode

Paged KV cache for SenseNova-U1 autoregressive decode, and a CUDA graph over it.

pipeline_sensenova_u1

SenseNova-U1 Pipeline for vLLM-Omni.

sensenova_u1_transformer

Qwen3 LLM with Mixture-of-Tokenizers (MoT) for SenseNova-U1.

SenseNovaU1Pipeline

Bases: Module, SupportsComponentDiscovery, DiffusionPipelineProfilerMixin, CFGParallelMixin, LoraLoaderMixin

SenseNova-U1 text-to-image and image-to-image pipeline for vllm-omni.

Builds the full model graph internally: - language_model: SenseNovaU1ForCausalLM (TP-aware) - vision_model: NEOVisionModel (understanding branch) - fm_modules: ModuleDict with vision_model_mot_gen, timestep_embedder, fm_head, etc.

img2img (image editing) is triggered when multi_modal_data["image"] is present in the prompt dict. The pipeline then uses triple KV caches (condition / img_condition / uncondition) with dual CFG (cfg_scale + img_cfg_scale).

denoising_transformer instance-attribute

denoising_transformer = SenseNovaU1DenoisingAdapter(
    self.language_model
)

device instance-attribute

device = get_local_device()

downsample_ratio instance-attribute

downsample_ratio = self.model_cfg.downsample_ratio

fm_modules instance-attribute

fm_modules = nn.ModuleDict(
    {
        "vision_model_mot_gen": vision_model_mot_gen,
        "timestep_embedder": timestep_embedder,
        "fm_head": fm_head,
    }
)

img_context_token_id instance-attribute

img_context_token_id = self.tokenizer.convert_tokens_to_ids(
    IMG_CONTEXT_TOKEN
)

img_start_token_id instance-attribute

img_start_token_id = self.tokenizer.convert_tokens_to_ids(
    IMG_START_TOKEN
)

language_model instance-attribute

language_model = SenseNovaU1ForCausalLM(
    self.llm_cfg, prefix="language_model"
)

llm_cfg instance-attribute

llm_cfg = self.model_cfg.llm_config

local_model_path instance-attribute

local_model_path = _resolve_model_path(model_path)

merge_size instance-attribute

merge_size = merge_size

model_cfg instance-attribute

model_cfg = SenseNovaU1Config.from_pretrained(
    self.local_model_path
)

od_config instance-attribute

od_config = od_config

patch_size instance-attribute

patch_size = patch_size

stacked_params_mapping class-attribute

stacked_params_mapping: list[tuple[str, str, str | int]] = [
    (".qkv_proj_mot_gen", ".q_proj_mot_gen", "q"),
    (".qkv_proj_mot_gen", ".k_proj_mot_gen", "k"),
    (".qkv_proj_mot_gen", ".v_proj_mot_gen", "v"),
    (".qkv_proj", ".q_proj", "q"),
    (".qkv_proj", ".k_proj", "k"),
    (".qkv_proj", ".v_proj", "v"),
    (".gate_up_proj", ".gate_proj", 0),
    (".gate_up_proj", ".up_proj", 1),
]

support_image_input class-attribute instance-attribute

support_image_input = True

tokenizer instance-attribute

tokenizer = AutoTokenizer.from_pretrained(
    self.local_model_path
)

transformer instance-attribute

transformer = self.language_model.model

vis_cfg instance-attribute

vis_cfg = self.model_cfg.vision_config

vision_model instance-attribute

vision_model = NEOVisionModel(self.vis_cfg)

weights_sources instance-attribute

weights_sources = [
    DiffusersPipelineLoader.ComponentSource(
        model_or_path=self.local_model_path,
        subfolder=None,
        revision=od_config.revision,
        prefix="",
        fall_back_to_pt=False,
    )
]

combine_cfg_noise

combine_cfg_noise(
    out_cond,
    out_uncond,
    cfg_scale,
    cfg_norm,
    kwargs: dict[str, Any] | None = None,
)

combine_multi_branch_cfg_noise

combine_multi_branch_cfg_noise(
    predictions, true_cfg_scale, cfg_normalize
)

forward

load_lora_weights

load_lora_weights(
    pretrained_model_name_or_path: str | list[str],
    adapter_name: str | None = None,
) -> None

Fuse a distilled few-step LoRA into the weights.

The checkpoints ship kohya lora_down/lora_up/alpha names, renamed here to the Diffusers names before load_lora_into_module routes each delta into its slice of the fused projections.

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

predict_noise

predict_noise(**kwargs)

release_captured_graphs

release_captured_graphs() -> None

Drop the reused paged cache and the graphs captured against it.

Sleep level 2 discards the memory a capture recorded, so anything held across requests has to go with it. The next request rebuilds both.

get_sensenova_u1_post_process_func

get_sensenova_u1_post_process_func(
    od_config: OmniDiffusionConfig,
)