Skip to content

vllm_omni.diffusion.models.interface

DecodedChunkConsumer module-attribute

DecodedChunkConsumer = Callable[['torch.Tensor'], None]

ReferenceVideoDecodeSpec dataclass

keep class-attribute instance-attribute

keep: Literal['first', 'last'] = 'first'

max_frames class-attribute instance-attribute

max_frames: int | None = None

SupportAudioInput

Bases: Protocol

support_audio_input class-attribute

support_audio_input: bool = True

SupportAudioOutput

Bases: Protocol

support_audio_output class-attribute

support_audio_output: bool = True

SupportImageInput

Bases: Protocol

color_format class-attribute

color_format: str = 'RGB'

support_image_input class-attribute

support_image_input: bool = True

SupportsChunkedVAEDecode

Bases: Protocol

Optional capability for VAEs that stream decoded temporal chunks.

Implementations must drain decoding before surfacing callback errors. In a distributed VAE, every rank invokes the method so collectives stay in lockstep, while on_chunk is called only on the rank that owns output.

chunk_value_range is the closed interval the delivered floats occupy, which differs per checkpoint: a Wan VAE emits (-1.0, 1.0) while a MiniMax-H3 VAE reverts through its processor to (0.0, 1.0). A consumer cannot quantize chunks to pixels without it, so it is part of the capability rather than knowledge each consumer hard-codes per model.

chunk_value_range class-attribute

chunk_value_range: tuple[float, float]

decode_with_chunks

decode_with_chunks(
    z: Tensor, *, on_chunk: DecodedChunkConsumer
) -> None

Decode z and synchronously deliver committed chunks in order.

SupportsComponentDiscovery

Bases: Protocol

Declares which submodules serve as pipeline components.

Used by the framework to locate DiT, encoder, and VAE modules for CPU offload, HSDP sharding, and other operations that need to know the pipeline's internal structure.

All attribute names support dotted paths for nested submodules (e.g. "pipe.transformer").

Attributes:

Name Type Description
_dit_modules list[str]

Denoising submodules (on GPU during diffusion).

_encoder_modules list[str]

Encoder submodules (offloaded during diffusion).

_vae_modules list[str]

VAE(s) (always on GPU).

_resident_modules list[str]

Extra modules pinned on GPU during layerwise offloading. Optional, defaults to [].

SupportsInteractionApply

Bases: Protocol

Optional protocol for pipelines with unified mid-generation, chunk-boundary hooks.

apply_interaction_at_chunk_boundary

apply_interaction_at_chunk_boundary(
    state: StepRequestState,
) -> None

Advance queued interactions before the next generation chunk.

peek_chunk_media

peek_chunk_media(state: StepRequestState) -> ChunkMediaSpec

Return the media timeline and latent frame count for the upcoming/current chunk.

Useful when interaction handler needs interpolation/integration on a frame-by-frame basis, and/or for backpressure/pacing.

prepare_next_chunk

prepare_next_chunk(state: StepRequestState) -> None

Set up pipeline state for the next chunk after interaction apply.

Default implementations is a no-op; model-specific pipelines override when chunk transitions require latent/history bookkeeping.

SupportsRequestScopedCacheDiT

Bases: Protocol

Optional protocol for pipelines that own Cache-DiT hook transitions.

adopt_cache_dit_backend

adopt_cache_dit_backend(backend: CacheDiTBackend) -> None

Assume ownership of an enabled Cache-DiT backend.

is_cache_dit_enabled

is_cache_dit_enabled() -> bool

Return whether this pipeline currently has Cache-DiT installed.

SupportsStepExecution

Bases: Protocol

State-driven step-level execution protocol for diffusion pipelines.

Pipelines should split request-level forward() into: prepare_encode() (one-time request setup), denoise_step() (one denoise forward), step_scheduler() (one scheduler update), and post_decode() (final decode).

A pipeline may additionally set the optional class attribute supports_chunk_step_grouping = True to declare that its request state may advance through every denoise step of its current chunk without a serving-scheduler cycle in between (a capability, not a policy: the runner that owns the scheduling decision chooses whether to group steps). It is optional so it does not become part of the runtime protocol check.

supports_step_execution class-attribute

supports_step_execution: bool = True

denoise_step

denoise_step(
    input_batch: InputBatch,
    *,
    states: Sequence[StepRequestState] | None = None,
    **kwargs: Any,
) -> Tensor | None

Run one denoise forward on the runner-assembled batch.

post_decode

post_decode(
    state: StepRequestState, **kwargs: Any
) -> DiffusionOutput

Decode output after denoise loop or at a partial chunk boundary.

prepare_encode

prepare_encode(
    state: StepRequestState, **kwargs: Any
) -> StepRequestState

Prepare request-level inputs and return initialized state.

step_scheduler

step_scheduler(
    state: StepRequestState,
    noise_pred: Tensor,
    **kwargs: Any,
) -> None

Run one scheduler step.

adopt_request_scoped_cache_dit

adopt_request_scoped_cache_dit(
    pipeline: object, backend: CacheDiTBackend
) -> bool

Transfer an enabled Cache-DiT backend to an opted-in pipeline.

is_request_scoped_cache_dit_enabled

is_request_scoped_cache_dit_enabled(
    pipeline: object,
) -> bool

Read Cache-DiT state from a pipeline that owns its lifecycle.

supports_chunked_vae_decode

supports_chunked_vae_decode(
    vae: object,
) -> TypeGuard[SupportsChunkedVAEDecode]

Return whether vae exposes the optional chunked decode capability.

supports_interaction_apply

supports_interaction_apply(pipeline: object) -> bool

Return whether pipeline implements :class:SupportsInteractionApply.

supports_step_execution

supports_step_execution(pipeline: object) -> bool

Return whether pipeline implements :class:SupportsStepExecution.