Skip to content

vllm_omni.diffusion.models.lingbot_world.dmd_block

One AR block of LingBot World's causal DMD sampler.

The pipeline drives this in two ways -- request mode loops generate_block over every block, stepwise execution calls probe_step / apply_transition one denoise step at a time and commit_block_kv from post_decode -- so the math lives here once and cannot drift between the two modes.

By default a block is four probes followed by one clean-x0 KV commit. The experimental last-step reuse mode instead commits the fourth noisy probe. Whether a call sees paged (session-bound) or request-local KV is decided by the ar argument, so this class never reaches back into the pipeline.

logger module-attribute

logger = init_logger(__name__)

ARBlockContext dataclass

The bound AR-Diffusion session one block runs against.

Built by the pipeline from the state the runner bound for this invocation; None in place of a context means request-local contiguous KV.

branch instance-attribute

branch: str

cross_attention instance-attribute

cross_attention: list[LingBotAttentionCache]

state instance-attribute

state: ARDiffusionKVState

LingBotDMDBlockRunner

Runs the DMD probes and the KV commit for one latent-frame block.

device instance-attribute

device = device

enforce_eager instance-attribute

enforce_eager = enforce_eager

reuse_last_step_kv instance-attribute

reuse_last_step_kv = reuse_last_step_kv

transformer instance-attribute

transformer = transformer

apply_transition

apply_transition(
    current_latents: Tensor,
    flow_prediction: Tensor,
    sigma: float,
    *,
    next_sigma: float | None,
    generator: Generator,
) -> Tensor

Invert the flow to x0, then re-noise at next_sigma.

next_sigma=None marks the final step, so where the caller sits in its own loop stays out of this function.

commit_block_kv

commit_block_kv(
    *,
    latents: Tensor,
    condition: Tensor,
    camera: Tensor,
    prompt_embeds: Tensor,
    cache: LingBotTransformerCache | None,
    ar: ARBlockContext | None,
    start_frame: int,
    camera_cache: CameraModulationCache | None = None,
) -> None

Write the finished block's clean x0 into KV and commit its pages.

By default this is the fifth transformer call, at t=0. Experimental reuse only finalizes the pages written by the fourth probe. On the stepwise path both modes finalize from post_decode(). camera_cache is as in probe_step.

generate_block

generate_block(
    *,
    condition: Tensor,
    camera: Tensor,
    prompt_embeds: Tensor,
    cache: LingBotTransformerCache | None,
    ar: ARBlockContext | None,
    start_frame: int,
    schedule: tuple[tuple[float, float], ...],
    generator: Generator,
    progress_bar: TqdmProgressBar[Any],
) -> Tensor

Request mode: all probes of one block, then its commit.

probe_step

probe_step(
    *,
    current_latents: Tensor,
    condition: Tensor,
    camera: Tensor,
    prompt_embeds: Tensor,
    cache: LingBotTransformerCache | None,
    ar: ARBlockContext | None,
    start_frame: int,
    timestep_value: float,
    step_index: int,
    camera_cache: CameraModulationCache | None = None,
) -> Tensor

Predict flow for one denoise step.

Normally only clean x0 enters KV. The experimental reuse mode commits the final noisy probe instead; the fixed DMD schedule has four probes.

camera_cache is the block's camera-modulation cache when the caller holds it across separate calls (the stepwise path); generate_block opens the window itself instead.