vllm_omni.diffusion.models.lingbot_world.dmd_block ¶
One AR block of LingBot World's causal DMD sampler.
The pipeline drives this in two ways -- request mode loops generate_block over every block, stepwise execution calls probe_step / apply_transition one denoise step at a time and commit_block_kv from post_decode -- so the math lives here once and cannot drift between the two modes.
By default a block is four probes followed by one clean-x0 KV commit. The experimental last-step reuse mode instead commits the fourth noisy probe. Whether a call sees paged (session-bound) or request-local KV is decided by the ar argument, so this class never reaches back into the pipeline.
ARBlockContext dataclass ¶
The bound AR-Diffusion session one block runs against.
Built by the pipeline from the state the runner bound for this invocation; None in place of a context means request-local contiguous KV.
LingBotDMDBlockRunner ¶
Runs the DMD probes and the KV commit for one latent-frame block.
apply_transition ¶
apply_transition(
current_latents: Tensor,
flow_prediction: Tensor,
sigma: float,
*,
next_sigma: float | None,
generator: Generator,
) -> Tensor
Invert the flow to x0, then re-noise at next_sigma.
next_sigma=None marks the final step, so where the caller sits in its own loop stays out of this function.
commit_block_kv ¶
commit_block_kv(
*,
latents: Tensor,
condition: Tensor,
camera: Tensor,
prompt_embeds: Tensor,
cache: LingBotTransformerCache | None,
ar: ARBlockContext | None,
start_frame: int,
camera_cache: CameraModulationCache | None = None,
) -> None
Write the finished block's clean x0 into KV and commit its pages.
By default this is the fifth transformer call, at t=0. Experimental reuse only finalizes the pages written by the fourth probe. On the stepwise path both modes finalize from post_decode(). camera_cache is as in probe_step.
generate_block ¶
generate_block(
*,
condition: Tensor,
camera: Tensor,
prompt_embeds: Tensor,
cache: LingBotTransformerCache | None,
ar: ARBlockContext | None,
start_frame: int,
schedule: tuple[tuple[float, float], ...],
generator: Generator,
progress_bar: TqdmProgressBar[Any],
) -> Tensor
Request mode: all probes of one block, then its commit.
probe_step ¶
probe_step(
*,
current_latents: Tensor,
condition: Tensor,
camera: Tensor,
prompt_embeds: Tensor,
cache: LingBotTransformerCache | None,
ar: ARBlockContext | None,
start_frame: int,
timestep_value: float,
step_index: int,
camera_cache: CameraModulationCache | None = None,
) -> Tensor
Predict flow for one denoise step.
Normally only clean x0 enters KV. The experimental reuse mode commits the final noisy probe instead; the fixed DMD schedule has four probes.
camera_cache is the block's camera-modulation cache when the caller holds it across separate calls (the stepwise path); generate_block opens the window itself instead.