vllm_omni.diffusion.offloader.layerwise_backend ¶
LayerWiseOffloadBackend ¶
Bases: OffloadBackend
Layer-wise (block-level) offloading backend.
Implements sliding window offloading where only a small number of transformer blocks reside on GPU at a time. Blocks are prefetched asynchronously while previous blocks compute, and freed after use.
LayerwiseOffloadHook ¶
Bases: ModelHook
Hook for layerwise (transformer-block-wise) CPU offloading.
The hook instance retains parameters for both the current registered block module and those for the next block, as well as flattened CPU tensors which record the parameters of the current block module, so that these parameters could be re-materialized on device in an overlapping way. This hook should be registered to each of the transformer blocks in DiT module(s) of the target pipeline.
Based on implementations from: https://github.com/sgl-project/sglang/blob/v0.5.8/python/sglang/multimodal_gen/runtime/utils/layerwise_offload.py
dtype_cpu_flattened_weights instance-attribute ¶
dtype_cpu_flattened_weights: dict[dtype, Tensor] = {}
is_materialized property ¶
is_materialized: bool
Check whether this block's parameters hold real data on device.
offload_layer ¶
Free GPU memory for layer by replacing tensors with empty placeholders. This function does not actually offload weights from GPU back to CPU.
prefetch_layer ¶
prefetch_layer(non_blocking: bool = True) -> None
Copy layer weights from CPU -> GPU.
Pre-fetch target block in an asynchronous way with compute - memory copy overlap, with non_blocking set to True.
restore_next_block ¶
Detach the next block from this hook's host backing store.
apply_block_hook ¶
apply_block_hook(
module: Module,
next_block: Module,
device: device,
stream: Stream | None = None,
pin_memory: bool = True,
*,
materialization_probe_tensor: Tensor | None = None,
) -> LayerwiseOffloadHook
disable_plan_encoder_layerwise_offload ¶
disable_plan_encoder_layerwise_offload(
module: Module, *, restore_weights: bool = True
) -> None
Remove hooks installed by :func:enable_plan_encoder_layerwise_offload.
enable_plan_encoder_layerwise_offload ¶
enable_plan_encoder_layerwise_offload(
component: ResolvedComponent,
*,
device: device,
stream: Stream,
pin_memory: bool,
) -> bool
Apply rank-local layerwise hooks to the encoder's resolved stacks.