vllm_omni.diffusion.models.lingbot_world ¶
Modules:
| Name | Description |
|---|---|
actions | Realtime keyboard controls for LingBot World 2.0. |
camera | Camera trajectory loading and geometry for LingBot World v2 conditioning. |
dmd_block | One AR block of LingBot World's causal DMD sampler. |
pipeline | Request-scoped LingBot-World v2 causal DMD pipeline. |
transformer | Checkpoint-compatible causal DiT for the LingBot World v2 model package. |
utils | Configuration helpers for the LingBot World pipeline. |
CausalLingBotWorldTransformer3DModel ¶
Bases: Module
Checkpoint-compatible causal LingBot World video transformer.
blocks instance-attribute ¶
blocks = nn.ModuleList(
[
LingBotAttentionBlock(
dim,
num_attention_heads,
ffn_dim=ffn_dim,
cross_attn_norm=cross_attn_norm,
eps=eps,
quant_config=quant_config,
prefix=_projection_prefix(
prefix, f"blocks.{index}"
),
)
for index in range(num_layers)
]
)
c2ws_hidden_states_layer1 instance-attribute ¶
c2ws_hidden_states_layer1 = ColumnParallelLinear(
dim,
dim,
bias=True,
gather_output=False,
return_bias=False,
quant_config=quant_config,
prefix=_projection_prefix(
prefix, "c2ws_hidden_states_layer1"
),
)
c2ws_hidden_states_layer2 instance-attribute ¶
c2ws_hidden_states_layer2 = RowParallelLinear(
dim,
dim,
bias=True,
input_is_parallel=True,
return_bias=False,
quant_config=quant_config,
prefix=_projection_prefix(
prefix, "c2ws_hidden_states_layer2"
),
)
config instance-attribute ¶
config = SimpleNamespace(
patch_size=patch_size,
num_attention_heads=num_attention_heads,
attention_head_dim=attention_head_dim,
in_channels=in_channels,
out_channels=out_channels,
text_dim=text_dim,
freq_dim=freq_dim,
ffn_dim=ffn_dim,
num_layers=num_layers,
cross_attn_norm=cross_attn_norm,
eps=eps,
image_dim=image_dim,
added_kv_proj_dim=added_kv_proj_dim,
rope_max_seq_len=rope_max_seq_len,
pos_embed_seq_len=pos_embed_seq_len,
qk_norm=qk_norm,
sink_size=sink_size,
num_frames_per_block=num_frames_per_block,
sliding_window_num_frames=sliding_window_num_frames,
local_attn_size=local_attn_size,
)
packed_modules_mapping class-attribute instance-attribute ¶
patch_embedding instance-attribute ¶
patch_embedding = Conv3dLayer(
in_channels=in_channels,
out_channels=dim,
kernel_size=patch_size,
stride=patch_size,
)
patch_embedding_wancamctrl instance-attribute ¶
text_embedding instance-attribute ¶
text_embedding = nn.Sequential(
nn.Linear(text_dim, dim),
nn.GELU(approximate="tanh"),
nn.Linear(dim, dim),
)
time_embedding instance-attribute ¶
time_projection instance-attribute ¶
allocate_cache ¶
allocate_cache(
*,
batch_size: int,
latent_height: int,
latent_width: int,
device: device,
dtype: dtype,
) -> LingBotTransformerCache
Allocate one caller-owned cache using this transformer's geometry.
forward ¶
forward(
hidden_states: Tensor,
timestep: Tensor,
encoder_hidden_states: Tensor,
camera_hidden_states: Tensor,
*,
cache: LingBotTransformerCache,
start_frame: int,
update_cache: bool,
camera_modulation_cache: CameraModulationCache
| None = None,
) -> Tensor
LingBotCameraControlReducer ¶
Sample live state/script controls into one LingBot AR block.
State mode preserves held keys across chunks and guarantees a one-frame pulse when a press and release both arrive before the next chunk. Script mode is a finite FIFO padded with neutral frames after exhaustion.
LingBotWorldCausalDMDPipeline ¶
Bases: Module, SupportImageInput, SupportsComponentDiscovery, SupportsStepExecution, InteractionMixin, ProgressBarMixin, DiffusionPipelineProfilerMixin
LingBot-World v2 I2V generation with a request-local causal cache.
text_encoder instance-attribute ¶
text_encoder = from_pretrained_with_prefetch(
UMT5EncoderModel.from_pretrained,
model,
subfolder="text_encoder",
prefetch_list=subfolders,
local_files_only=local_files_only,
torch_dtype=dtype,
)
tokenizer instance-attribute ¶
tokenizer = from_pretrained_with_prefetch(
AutoTokenizer.from_pretrained,
model,
subfolder="tokenizer",
prefetch_list=subfolders,
local_files_only=local_files_only,
)
transformer instance-attribute ¶
transformer = (
CausalLingBotWorldTransformer3DModel.from_config(
transformer_config,
quant_config=getattr(
od_config, "quantization_config", None
),
prefix="transformer",
)
)
vae instance-attribute ¶
vae = from_pretrained_with_prefetch(
DistributedAutoencoderKLWan.from_pretrained,
model,
subfolder="vae",
prefetch_list=subfolders,
local_files_only=local_files_only,
torch_dtype=dtype,
)
vae_scale_factor_spatial instance-attribute ¶
vae_scale_factor_temporal instance-attribute ¶
weights_sources instance-attribute ¶
weights_sources = [
DiffusersPipelineLoader.ComponentSource(
model_or_path=model,
subfolder="transformer",
revision=None,
prefix="transformer.",
fall_back_to_pt=True,
)
]
ar_diffusion_kv_cache_spec ¶
Describe the fixed worker-local cache geometry for realtime ticks.
bind_ar_diffusion_state ¶
denoise_step ¶
denoise_step(
input_batch: InputBatch,
*,
states: Sequence[StepRequestState] | None = None,
**kwargs: Any,
) -> Tensor | None
encode_prompt ¶
Encode prompt with UMT5, reusing the result for a prompt already encoded.
The encode depends only on the whitespace-normalised text, the sequence length and the dtype -- the tokenizer and text encoder are fixed once loaded -- so those three are the key. That covers every caller the same way: a realtime tick repeats its session's prompt on every block, a stepwise request encodes once per request, and a server tends to open many sessions with the same scene prompt.
A hit returns the stored tensor itself. Callers only read it: the cross-attention projection and the DMD transformer both work out of place, so a shared encode is never changed under another session.
peek_chunk_media ¶
peek_chunk_media(state: StepRequestState) -> ChunkMediaSpec
Expose this chunk's decoded media extent and latent step count.
prepare_next_chunk ¶
prepare_next_chunk(state: StepRequestState) -> None
Prepare the next AR block after chunk-boundary interaction apply.
step_scheduler ¶
step_scheduler(
state: StepRequestState,
noise_pred: Tensor,
**kwargs: Any,
) -> None
get_lingbot_world_post_process_func ¶
get_lingbot_world_post_process_func(
od_config: OmniDiffusionConfig,
) -> Callable[..., Any]
get_lingbot_world_pre_process_func ¶
get_lingbot_world_pre_process_func(
od_config: OmniDiffusionConfig,
) -> Callable[[OmniDiffusionRequest], OmniDiffusionRequest]
Materialize request-local files once before dispatching GPU workers.