Skip to content

vllm_omni.diffusion.models.lingbot_world

Modules:

Name Description
actions

Realtime keyboard controls for LingBot World 2.0.

camera

Camera trajectory loading and geometry for LingBot World v2 conditioning.

dmd_block

One AR block of LingBot World's causal DMD sampler.

pipeline

Request-scoped LingBot-World v2 causal DMD pipeline.

transformer

Checkpoint-compatible causal DiT for the LingBot World v2 model package.

utils

Configuration helpers for the LingBot World pipeline.

CausalLingBotWorldTransformer3DModel

Bases: Module

Checkpoint-compatible causal LingBot World video transformer.

blocks instance-attribute

blocks = nn.ModuleList(
    [
        LingBotAttentionBlock(
            dim,
            num_attention_heads,
            ffn_dim=ffn_dim,
            cross_attn_norm=cross_attn_norm,
            eps=eps,
            quant_config=quant_config,
            prefix=_projection_prefix(
                prefix, f"blocks.{index}"
            ),
        )
        for index in range(num_layers)
    ]
)

c2ws_hidden_states_layer1 instance-attribute

c2ws_hidden_states_layer1 = ColumnParallelLinear(
    dim,
    dim,
    bias=True,
    gather_output=False,
    return_bias=False,
    quant_config=quant_config,
    prefix=_projection_prefix(
        prefix, "c2ws_hidden_states_layer1"
    ),
)

c2ws_hidden_states_layer2 instance-attribute

c2ws_hidden_states_layer2 = RowParallelLinear(
    dim,
    dim,
    bias=True,
    input_is_parallel=True,
    return_bias=False,
    quant_config=quant_config,
    prefix=_projection_prefix(
        prefix, "c2ws_hidden_states_layer2"
    ),
)

config instance-attribute

config = SimpleNamespace(
    patch_size=patch_size,
    num_attention_heads=num_attention_heads,
    attention_head_dim=attention_head_dim,
    in_channels=in_channels,
    out_channels=out_channels,
    text_dim=text_dim,
    freq_dim=freq_dim,
    ffn_dim=ffn_dim,
    num_layers=num_layers,
    cross_attn_norm=cross_attn_norm,
    eps=eps,
    image_dim=image_dim,
    added_kv_proj_dim=added_kv_proj_dim,
    rope_max_seq_len=rope_max_seq_len,
    pos_embed_seq_len=pos_embed_seq_len,
    qk_norm=qk_norm,
    sink_size=sink_size,
    num_frames_per_block=num_frames_per_block,
    sliding_window_num_frames=sliding_window_num_frames,
    local_attn_size=local_attn_size,
)

dim instance-attribute

dim = dim

dtype property

dtype: dtype

Return the dtype used by the transformer parameters.

head instance-attribute

head = _LingBotHead(dim, out_channels, patch_size, eps)

packed_modules_mapping class-attribute instance-attribute

packed_modules_mapping = {'qkv': ['q', 'k', 'v']}

patch_embedding instance-attribute

patch_embedding = Conv3dLayer(
    in_channels=in_channels,
    out_channels=dim,
    kernel_size=patch_size,
    stride=patch_size,
)

patch_embedding_wancamctrl instance-attribute

patch_embedding_wancamctrl = _LingBotCameraPatchEmbedding(
    6 * 8 * 8, dim, patch_size
)

sp_output_gather instance-attribute

sp_output_gather = nn.Identity()

sp_prepare instance-attribute

sp_prepare = _LingBotSPPrepare()

text_embedding instance-attribute

text_embedding = nn.Sequential(
    nn.Linear(text_dim, dim),
    nn.GELU(approximate="tanh"),
    nn.Linear(dim, dim),
)

time_embedding instance-attribute

time_embedding = nn.Sequential(
    nn.Linear(freq_dim, dim),
    nn.SiLU(),
    nn.Linear(dim, dim),
)

time_projection instance-attribute

time_projection = nn.Sequential(
    nn.SiLU(), nn.Linear(dim, dim * 6)
)

allocate_cache

allocate_cache(
    *,
    batch_size: int,
    latent_height: int,
    latent_width: int,
    device: device,
    dtype: dtype,
) -> LingBotTransformerCache

Allocate one caller-owned cache using this transformer's geometry.

forward

forward(
    hidden_states: Tensor,
    timestep: Tensor,
    encoder_hidden_states: Tensor,
    camera_hidden_states: Tensor,
    *,
    cache: LingBotTransformerCache,
    start_frame: int,
    update_cache: bool,
    camera_modulation_cache: CameraModulationCache
    | None = None,
) -> Tensor

from_config classmethod

from_config(
    config: dict[str, Any],
    *,
    quant_config: QuantizationConfig | None = None,
    prefix: str = "",
) -> Self

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

Load exact checkpoint names, delegating TP sharding to parameters.

LingBotCameraControlReducer

Sample live state/script controls into one LingBot AR block.

State mode preserves held keys across chunks and guarantees a one-frame pulse when a press and release both arrive before the next chunk. Script mode is a finite FIFO padded with neutral frames after exhaustion.

commit

commit(prepared: ARDiffusionPreparedControls) -> None

prepare

prepare(
    *,
    current_controls: Mapping[str, ARDiffusionControlInput],
    events: Sequence[ARDiffusionSessionEvent],
    chunk_index: int,
) -> ARDiffusionPreparedControls

reset

reset() -> None

LingBotWorldCausalDMDPipeline

Bases: Module, SupportImageInput, SupportsComponentDiscovery, SupportsStepExecution, InteractionMixin, ProgressBarMixin, DiffusionPipelineProfilerMixin

LingBot-World v2 I2V generation with a request-local causal cache.

device instance-attribute

device = get_local_device()

dummy_run_num_frames class-attribute

dummy_run_num_frames: int = 0

od_config instance-attribute

od_config = od_config

scheduler instance-attribute

scheduler = FlowUniPCMultistepScheduler(**scheduler_kwargs)

supports_chunk_step_grouping class-attribute

supports_chunk_step_grouping: bool = True

supports_step_execution class-attribute

supports_step_execution: bool = True

text_encoder instance-attribute

text_encoder = from_pretrained_with_prefetch(
    UMT5EncoderModel.from_pretrained,
    model,
    subfolder="text_encoder",
    prefetch_list=subfolders,
    local_files_only=local_files_only,
    torch_dtype=dtype,
)

tokenizer instance-attribute

tokenizer = from_pretrained_with_prefetch(
    AutoTokenizer.from_pretrained,
    model,
    subfolder="tokenizer",
    prefetch_list=subfolders,
    local_files_only=local_files_only,
)

transformer instance-attribute

transformer = (
    CausalLingBotWorldTransformer3DModel.from_config(
        transformer_config,
        quant_config=getattr(
            od_config, "quantization_config", None
        ),
        prefix="transformer",
    )
)

vae instance-attribute

vae = from_pretrained_with_prefetch(
    DistributedAutoencoderKLWan.from_pretrained,
    model,
    subfolder="vae",
    prefetch_list=subfolders,
    local_files_only=local_files_only,
    torch_dtype=dtype,
)

vae_scale_factor_spatial instance-attribute

vae_scale_factor_spatial = int(
    getattr(self.vae.config, "scale_factor_spatial", 8)
)

vae_scale_factor_temporal instance-attribute

vae_scale_factor_temporal = int(
    getattr(self.vae.config, "scale_factor_temporal", 4)
)

weights_sources instance-attribute

weights_sources = [
    DiffusersPipelineLoader.ComponentSource(
        model_or_path=model,
        subfolder="transformer",
        revision=None,
        prefix="transformer.",
        fall_back_to_pt=True,
    )
]

ar_diffusion_kv_cache_spec

ar_diffusion_kv_cache_spec() -> ARDiffusionKVCacheSpec

Describe the fixed worker-local cache geometry for realtime ticks.

bind_ar_diffusion_state

bind_ar_diffusion_state(
    session_id: str, state: ARDiffusionKVState
) -> Iterator[None]

close_ar_diffusion_session

close_ar_diffusion_session(session_id: str) -> None

denoise_step

denoise_step(
    input_batch: InputBatch,
    *,
    states: Sequence[StepRequestState] | None = None,
    **kwargs: Any,
) -> Tensor | None

encode_prompt

encode_prompt(
    prompt: str, *, max_sequence_length: int, dtype: dtype
) -> Tensor

Encode prompt with UMT5, reusing the result for a prompt already encoded.

The encode depends only on the whitespace-normalised text, the sequence length and the dtype -- the tokenizer and text encoder are fixed once loaded -- so those three are the key. That covers every caller the same way: a realtime tick repeats its session's prompt on every block, a stepwise request encodes once per request, and a server tends to open many sessions with the same scene prompt.

A hit returns the stored tensor itself. Callers only read it: the cross-attention projection and the DMD transformer both work out of place, so a shared encode is never changed under another session.

forward

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

peek_chunk_media

peek_chunk_media(state: StepRequestState) -> ChunkMediaSpec

Expose this chunk's decoded media extent and latent step count.

post_decode

post_decode(
    state: StepRequestState, **kwargs: Any
) -> DiffusionOutput

prepare_encode

prepare_encode(
    state: StepRequestState, **kwargs: Any
) -> StepRequestState

prepare_next_chunk

prepare_next_chunk(state: StepRequestState) -> None

Prepare the next AR block after chunk-boundary interaction apply.

reset_ar_diffusion_session

reset_ar_diffusion_session(session_id: str) -> None

step_scheduler

step_scheduler(
    state: StepRequestState,
    noise_pred: Tensor,
    **kwargs: Any,
) -> None

get_lingbot_world_post_process_func

get_lingbot_world_post_process_func(
    od_config: OmniDiffusionConfig,
) -> Callable[..., Any]

get_lingbot_world_pre_process_func

get_lingbot_world_pre_process_func(
    od_config: OmniDiffusionConfig,
) -> Callable[[OmniDiffusionRequest], OmniDiffusionRequest]

Materialize request-local files once before dispatching GPU workers.