Skip to content

vllm_omni.diffusion.models.sana_video.pipeline_sana_video

ASPECT_RATIO_480_BIN module-attribute

ASPECT_RATIO_480_BIN = {
    "0.5": [448.0, 896.0],
    "0.57": [480.0, 832.0],
    "0.68": [528.0, 768.0],
    "0.78": [560.0, 720.0],
    "1.0": [624.0, 624.0],
    "1.13": [672.0, 592.0],
    "1.29": [720.0, 560.0],
    "1.46": [768.0, 528.0],
    "1.67": [816.0, 496.0],
    "1.75": [832.0, 480.0],
    "2.0": [896.0, 448.0],
}

ASPECT_RATIO_720_BIN module-attribute

ASPECT_RATIO_720_BIN = {
    "0.5": [672.0, 1344.0],
    "0.57": [704.0, 1280.0],
    "0.68": [800.0, 1152.0],
    "0.78": [832.0, 1088.0],
    "1.0": [960.0, 960.0],
    "1.13": [1024.0, 896.0],
    "1.29": [1088.0, 832.0],
    "1.46": [1152.0, 800.0],
    "1.67": [1248.0, 736.0],
    "1.75": [1280.0, 704.0],
    "2.0": [1344.0, 672.0],
}

SANA_VIDEO_COMPLEX_HUMAN_INSTRUCTION module-attribute

SANA_VIDEO_COMPLEX_HUMAN_INSTRUCTION = (
    "Given a user prompt, generate an 'Enhanced prompt' that provides detailed visual descriptions suitable for video generation. Evaluate the level of detail in the user prompt:",
    "- If the prompt is simple, focus on adding specifics about colors, shapes, sizes, textures, motion, and temporal relationships to create vivid and dynamic scenes.",
    "- If the prompt is already detailed, refine and enhance the existing details slightly without overcomplicating.",
    "Here are examples of how to transform or refine prompts:",
    "- User Prompt: A cat sleeping -> Enhanced: A small, fluffy white cat slowly settling into a curled position, peacefully falling asleep on a warm sunny windowsill, with gentle sunlight filtering through surrounding pots of blooming red flowers.",
    "- User Prompt: A busy city street -> Enhanced: A bustling city street scene at dusk, featuring glowing street lamps gradually lighting up, a diverse crowd of people in colorful clothing walking past, and a double-decker bus smoothly passing by towering glass skyscrapers.",
    "Please generate only the enhanced description for the prompt below and avoid including any additional commentary or evaluations:",
    "User Prompt: ",
)

SanaVideoPipeline

Bases: Module, CFGParallelMixin, ProgressBarMixin, DiffusionPipelineProfilerMixin, SupportsComponentDiscovery

vLLM-Omni pipeline for text-to-video generation using Sana.

Parameters:

Name Type Description Default
tokenizer [`GemmaTokenizer`] or [`GemmaTokenizerFast`]

The tokenizer used to tokenize the prompt.

None
text_encoder [`Gemma2PreTrainedModel`]

Text encoder model to encode the input prompts.

None
vae [`DistributedAutoencoderKLWan`] or [`DistributedAutoencoderKLLTX2Video`]

Variational Auto-Encoder (VAE) Model to encode and decode videos to and from latent representations.

None
transformer [`SanaVideoTransformer3DModel`]

Conditional Transformer to denoise the input latents.

None
scheduler [`DPMSolverMultistepScheduler`]

A scheduler to be used in combination with transformer to denoise the encoded video latents.

None

default_num_inference_steps class-attribute instance-attribute

default_num_inference_steps = 50

device instance-attribute

device = get_local_device()

do_classifier_free_guidance property

do_classifier_free_guidance

dummy_run_num_frames class-attribute instance-attribute

dummy_run_num_frames = 25

guidance_scale property

guidance_scale

interrupt property

interrupt

num_timesteps property

num_timesteps

od_config instance-attribute

od_config = od_config

scheduler instance-attribute

scheduler = scheduler

supports_step_execution class-attribute instance-attribute

supports_step_execution = False

text_encoder instance-attribute

text_encoder = text_encoder

tokenizer instance-attribute

tokenizer = tokenizer

transformer instance-attribute

transformer = transformer

vae instance-attribute

vae = vae

vae_scale_factor instance-attribute

vae_scale_factor = self.vae_scale_factor_spatial

vae_scale_factor_spatial instance-attribute

vae_scale_factor_spatial = (
    self.vae.config.spatial_compression_ratio
)

vae_scale_factor_temporal instance-attribute

vae_scale_factor_temporal = (
    self.vae.config.temporal_compression_ratio
)

video_processor instance-attribute

video_processor = VideoProcessor(
    vae_scale_factor=self.vae_scale_factor_spatial
)

weights_sources instance-attribute

weights_sources: list[ComponentSource] = []

check_inputs

check_inputs(
    prompt,
    height,
    width,
    negative_prompt=None,
    prompt_embeds=None,
    negative_prompt_embeds=None,
    prompt_attention_mask=None,
    negative_prompt_attention_mask=None,
)

combine_cfg_noise

combine_cfg_noise(
    positive_noise_pred,
    negative_noise_pred,
    true_cfg_scale,
    cfg_normalize=False,
    kwargs=None,
)

diffuse

diffuse(
    latents: Tensor,
    timesteps: Tensor,
    prompt_embeds: Tensor,
    prompt_attention_mask: Tensor,
    negative_prompt_embeds: Tensor | None,
    negative_prompt_attention_mask: Tensor | None,
    guidance_scale: float,
    extra_step_kwargs: dict,
    dtype: dtype,
    output_slice: int | None,
) -> Tensor

encode_prompt

encode_prompt(
    prompt: str | list[str],
    do_classifier_free_guidance: bool = True,
    negative_prompt: str = "",
    num_videos_per_prompt: int = 1,
    device: device | None = None,
    prompt_embeds: Tensor | None = None,
    negative_prompt_embeds: Tensor | None = None,
    prompt_attention_mask: Tensor | None = None,
    negative_prompt_attention_mask: Tensor | None = None,
    clean_caption: bool = False,
    max_sequence_length: int = 300,
    complex_human_instruction: list[str] | None = None,
)

Encodes the prompt into text encoder hidden states.

Parameters:

Name Type Description Default
prompt `str` or `list[str]`, *optional*

prompt to be encoded

required
negative_prompt `str` or `list[str]`, *optional*

The prompt not to guide the video generation. If not defined, one has to pass negative_prompt_embeds instead. Ignored when not using guidance (i.e., ignored if guidance_scale is less than 1). For PixArt-Alpha, this should be "".

''
do_classifier_free_guidance `bool`, *optional*, defaults to `True`

whether to use classifier free guidance or not

True
num_videos_per_prompt `int`, *optional*, defaults to 1

number of videos that should be generated per prompt

1
device device | None

(torch.device, optional): torch device to place the resulting embeddings on

None
prompt_embeds `torch.Tensor`, *optional*

Pre-generated text embeddings. Can be used to easily tweak text inputs, e.g. prompt weighting. If not provided, text embeddings will be generated from prompt input argument.

None
negative_prompt_embeds `torch.Tensor`, *optional*

Pre-generated negative text embeddings. For Sana, it's should be the embeddings of the "" string.

None
clean_caption `bool`, defaults to `False`

If True, the function will preprocess and clean the provided caption before encoding.

False
max_sequence_length `int`, defaults to 300

Maximum sequence length to use for the prompt.

300
complex_human_instruction `list[str]`, defaults to `complex_human_instruction`

If complex_human_instruction is not empty, the function will use the complex Human instruction for the prompt.

None

forward

load_weights

load_weights(
    weights: Iterable[tuple[str, Tensor]],
) -> set[str]

predict_noise

predict_noise(*args, **kwargs) -> Tensor

prepare_extra_step_kwargs

prepare_extra_step_kwargs(generator, eta)

prepare_latents

prepare_latents(
    batch_size: int,
    num_channels_latents: int = 16,
    height: int = 480,
    width: int = 832,
    num_frames: int = 81,
    dtype: dtype | None = None,
    device: device | None = None,
    generator: Generator | list[Generator] | None = None,
    latents: Tensor | None = None,
) -> Tensor

get_sana_video_default_resolution

get_sana_video_default_resolution(
    sample_size: int,
) -> tuple[int, int]

get_sana_video_post_process_func

get_sana_video_post_process_func(
    od_config: OmniDiffusionConfig,
)

resolve_sana_video_sample_size

resolve_sana_video_sample_size(
    od_config: OmniDiffusionConfig,
) -> int

retrieve_timesteps

retrieve_timesteps(
    scheduler,
    num_inference_steps: int | None = None,
    device: str | device | None = None,
    timesteps: list[int] | None = None,
    sigmas: list[float] | None = None,
    **kwargs,
) -> tuple[Tensor, int]

Calls the scheduler's set_timesteps method and retrieves timesteps from the scheduler after the call. Handles custom timesteps. Any kwargs will be supplied to scheduler.set_timesteps.

Parameters:

Name Type Description Default
scheduler `SchedulerMixin`

The scheduler to get timesteps from.

required
num_inference_steps `int`

The number of diffusion steps used when generating samples with a pre-trained model. If used, timesteps must be None.

None
device `str` or `torch.device`, *optional*

The device to which the timesteps should be moved to. If None, the timesteps are not moved.

None
timesteps `list[int]`, *optional*

Custom timesteps used to override the timestep spacing strategy of the scheduler. If timesteps is passed, num_inference_steps and sigmas must be None.

None
sigmas `list[float]`, *optional*

Custom sigmas used to override the timestep spacing strategy of the scheduler. If sigmas is passed, num_inference_steps and timesteps must be None.

None

Returns:

Type Description
Tensor

tuple[torch.Tensor, int]: A tuple where the first element is the timestep schedule from the scheduler and the

int

second element is the number of inference steps.