vllm_omni.diffusion.offloader.plan_resolver ¶
One topology resolver shared by every diffusion offload backend.
Model topology is declared through several mechanisms that grew independently: :class:~vllm_omni.diffusion.models.interface.SupportsComponentDiscovery class variables, the pipeline-level :class:OffloadPlan, per-DiT _layerwise_offload_blocks_attrs, and a leaf-name rule for encoders. This module runs them in one defined order and returns one backend-neutral artifact.
The resolver is pure: it moves no tensor, installs no hook, writes no module attribute, and touches no process group. Everything it can reject is rejected here, so plan-dependent failures happen before a backend mutates the model.
BlockStack dataclass ¶
One ring of repeated blocks that a backend streams together.
attrs names the owner attributes that contributed blocks; backends use it to skip those children when they place the non-streamed remainder. It is filled for DiT stacks only — encoder stacks are hooked as whole stacks and their remainder is placed by tensor identity.
ResolvedComponent dataclass ¶
One pipeline component with its resolved, validated offload topology.
selected means the active selector covers this component. on_demand means the pipeline owns its residency through load_to_device / offload_to_cpu; a VAE staged by a legacy model plan is on_demand without being selectable through the public component grammar.
ResolvedOffloadPlan dataclass ¶
Backend-neutral topology for one pipeline under one offload config.
resolve_offload_plan ¶
resolve_offload_plan(
pipeline: Module, config: OffloadConfig
) -> ResolvedOffloadPlan
Resolve and validate one pipeline's offload topology.
Raises ValueError for every topology that the requested configuration cannot serve, before the caller places a module or installs a hook.