vllm_omni.diffusion.models.lingbot_video ¶
Modules:
| Name | Description |
|---|---|
image_condition | |
lingbot_video_transformer | |
pipeline_lingbot_video | |
request_utils | |
LingBotGenerationMode ¶
LingBotImageCondition dataclass ¶
LingBotRequestConfig dataclass ¶
LingBotVideoPipeline ¶
Bases: Module, SupportImageInput, ProgressBarMixin, SupportsComponentDiscovery
Native vLLM-Omni entry for LingBot-Video checkpoints.
The in-tree transformer supports both dense MLP blocks and routed MoE blocks. Fused expert kernels and the optional refiner/ transformer are not loaded or executed. t_thresh only selects a low-noise sigma schedule for the primary transformer; it does not enable automatic refiner orchestration.
default_image_negative_prompt instance-attribute ¶
default_image_negative_prompt = (
DEFAULT_NEGATIVE_PROMPT_IMAGE
)
processor instance-attribute ¶
processor = Qwen3VLProcessor.from_pretrained(
model,
subfolder=processor_subfolder,
local_files_only=local_files_only,
)
scheduler instance-attribute ¶
scheduler = FlowUniPCMultistepScheduler.from_pretrained(
model,
subfolder=scheduler_subfolder,
local_files_only=local_files_only,
)
text_encoder instance-attribute ¶
text_encoder = (
Qwen3VLForConditionalGeneration.from_pretrained(
model,
subfolder=text_encoder_subfolder,
**text_encoder_kwargs,
).to(self.device)
)
transformer instance-attribute ¶
transformer = (
LingBotVideoTransformer3DModel.from_pretrained(
model,
subfolder=transformer_subfolder,
torch_dtype=transformer_dtype,
local_files_only=local_files_only,
).to(self.device)
)
vae instance-attribute ¶
vae = AutoencoderKLWan.from_pretrained(
model,
subfolder=vae_subfolder,
torch_dtype=vae_dtype,
local_files_only=local_files_only,
).to(self.device)
apply_text_to_template staticmethod ¶
apply_text_to_template(
text: str, template: str = PROMPT_TEMPLATE
) -> str
encode_prompt ¶
encode_prompt(
prompt: str | list[str],
*,
images: Any | None = None,
device: str | device | None = None,
) -> tuple[Tensor, Tensor]
prepare_latents ¶
prepare_latents(
num_frames: int,
height: int,
width: int,
generator: Generator | None,
latents: Tensor | None,
device: device,
) -> Tensor
prepare_ti2v_image_condition ¶
prepare_ti2v_image_condition(
image: Image,
*,
height: int,
width: int,
generator: Generator | None = None,
) -> LingBotImageCondition
LingBotVideoTransformer3DModel ¶
Bases: ModelMixin, ConfigMixin
blocks instance-attribute ¶
blocks = nn.ModuleList(
[
LingBotVideoBlock(
hidden_size=hidden_size,
num_attention_heads=num_attention_heads,
intermediate_size=intermediate_size,
norm_eps=norm_eps,
qkv_bias=qkv_bias,
out_bias=out_bias,
num_experts=num_experts,
num_experts_per_tok=num_experts_per_tok,
moe_intermediate_size=moe_intermediate_size,
decoder_sparse_step=decoder_sparse_step,
mlp_only_layers=mlp_only_layers,
n_shared_experts=n_shared_experts,
score_func=score_func,
norm_topk_prob=norm_topk_prob,
n_group=n_group,
topk_group=topk_group,
routed_scaling_factor=routed_scaling_factor,
layer_idx=i,
)
for i in range(depth)
]
)
norm_out instance-attribute ¶
norm_out_modulation instance-attribute ¶
patch_embedder instance-attribute ¶
patch_embedder = nn.Linear(
in_channels * math.prod(patch_size),
hidden_size,
bias=patch_embed_bias,
)
proj_out instance-attribute ¶
text_embedder instance-attribute ¶
text_embedder = LingBotVideoTextEmbedder(
text_dim, hidden_size
)
time_embedder instance-attribute ¶
time_embedder = TimestepEmbedding(
freq_dim,
hidden_size,
act_fn="silu",
sample_proj_bias=timestep_mlp_bias,
)
time_modulation instance-attribute ¶
time_proj instance-attribute ¶
get_lingbot_video_post_process_func ¶
get_lingbot_video_post_process_func(
od_config: OmniDiffusionConfig,
)
get_lingbot_video_pre_process_func ¶
get_lingbot_video_pre_process_func(
od_config: OmniDiffusionConfig,
)
normalize_lingbot_num_frames ¶
Round a video length up to LingBot's causal VAE 4n+1 grid.
normalize_lingbot_request ¶
normalize_lingbot_request(
request: Any,
*,
default_negative_prompt: str,
default_image_negative_prompt: str,
default_height: int = 480,
default_width: int = 480,
default_num_frames: int = 81,
default_fps: int = 24,
default_num_inference_steps: int = 40,
default_guidance_scale: float = 6.0,
default_shift: float = 3.0,
default_output_type: str = "pt",
) -> LingBotRequestConfig
prepare_ti2v_image_condition ¶
prepare_ti2v_image_condition(
image: Image,
*,
height: int,
width: int,
vae: Module,
vision_patch_size: int,
device: device,
generator: Generator | None = None,
) -> LingBotImageCondition
resolve_lingbot_output_dimensions ¶
resolve_lingbot_output_dimensions(
*,
sampling_width: Any = None,
sampling_height: Any = None,
prompt_fields: Mapping[str, Any] | None = None,
extra_fields: Mapping[str, Any] | None = None,
default_width: int = 480,
default_height: int = 480,
) -> tuple[int, int]
Resolve the (width, height) that LingBot will use for generation.