vllm_omni.diffusion.models.sana_video.pipeline_sana_video ¶
ASPECT_RATIO_480_BIN module-attribute ¶
ASPECT_RATIO_480_BIN = {
"0.5": [448.0, 896.0],
"0.57": [480.0, 832.0],
"0.68": [528.0, 768.0],
"0.78": [560.0, 720.0],
"1.0": [624.0, 624.0],
"1.13": [672.0, 592.0],
"1.29": [720.0, 560.0],
"1.46": [768.0, 528.0],
"1.67": [816.0, 496.0],
"1.75": [832.0, 480.0],
"2.0": [896.0, 448.0],
}
ASPECT_RATIO_720_BIN module-attribute ¶
ASPECT_RATIO_720_BIN = {
"0.5": [672.0, 1344.0],
"0.57": [704.0, 1280.0],
"0.68": [800.0, 1152.0],
"0.78": [832.0, 1088.0],
"1.0": [960.0, 960.0],
"1.13": [1024.0, 896.0],
"1.29": [1088.0, 832.0],
"1.46": [1152.0, 800.0],
"1.67": [1248.0, 736.0],
"1.75": [1280.0, 704.0],
"2.0": [1344.0, 672.0],
}
SANA_VIDEO_COMPLEX_HUMAN_INSTRUCTION module-attribute ¶
SANA_VIDEO_COMPLEX_HUMAN_INSTRUCTION = (
"Given a user prompt, generate an 'Enhanced prompt' that provides detailed visual descriptions suitable for video generation. Evaluate the level of detail in the user prompt:",
"- If the prompt is simple, focus on adding specifics about colors, shapes, sizes, textures, motion, and temporal relationships to create vivid and dynamic scenes.",
"- If the prompt is already detailed, refine and enhance the existing details slightly without overcomplicating.",
"Here are examples of how to transform or refine prompts:",
"- User Prompt: A cat sleeping -> Enhanced: A small, fluffy white cat slowly settling into a curled position, peacefully falling asleep on a warm sunny windowsill, with gentle sunlight filtering through surrounding pots of blooming red flowers.",
"- User Prompt: A busy city street -> Enhanced: A bustling city street scene at dusk, featuring glowing street lamps gradually lighting up, a diverse crowd of people in colorful clothing walking past, and a double-decker bus smoothly passing by towering glass skyscrapers.",
"Please generate only the enhanced description for the prompt below and avoid including any additional commentary or evaluations:",
"User Prompt: ",
)
SanaVideoPipeline ¶
Bases: Module, CFGParallelMixin, ProgressBarMixin, DiffusionPipelineProfilerMixin, SupportsComponentDiscovery
vLLM-Omni pipeline for text-to-video generation using Sana.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
tokenizer | [`GemmaTokenizer`] or [`GemmaTokenizerFast`] | The tokenizer used to tokenize the prompt. | None |
text_encoder | [`Gemma2PreTrainedModel`] | Text encoder model to encode the input prompts. | None |
vae | [`DistributedAutoencoderKLWan`] or [`DistributedAutoencoderKLLTX2Video`] | Variational Auto-Encoder (VAE) Model to encode and decode videos to and from latent representations. | None |
transformer | [`SanaVideoTransformer3DModel`] | Conditional Transformer to denoise the input latents. | None |
scheduler | [`DPMSolverMultistepScheduler`] | A scheduler to be used in combination with | None |
vae_scale_factor_spatial instance-attribute ¶
vae_scale_factor_temporal instance-attribute ¶
video_processor instance-attribute ¶
check_inputs ¶
check_inputs(
prompt,
height,
width,
negative_prompt=None,
prompt_embeds=None,
negative_prompt_embeds=None,
prompt_attention_mask=None,
negative_prompt_attention_mask=None,
)
combine_cfg_noise ¶
combine_cfg_noise(
positive_noise_pred,
negative_noise_pred,
true_cfg_scale,
cfg_normalize=False,
kwargs=None,
)
diffuse ¶
diffuse(
latents: Tensor,
timesteps: Tensor,
prompt_embeds: Tensor,
prompt_attention_mask: Tensor,
negative_prompt_embeds: Tensor | None,
negative_prompt_attention_mask: Tensor | None,
guidance_scale: float,
extra_step_kwargs: dict,
dtype: dtype,
output_slice: int | None,
) -> Tensor
encode_prompt ¶
encode_prompt(
prompt: str | list[str],
do_classifier_free_guidance: bool = True,
negative_prompt: str = "",
num_videos_per_prompt: int = 1,
device: device | None = None,
prompt_embeds: Tensor | None = None,
negative_prompt_embeds: Tensor | None = None,
prompt_attention_mask: Tensor | None = None,
negative_prompt_attention_mask: Tensor | None = None,
clean_caption: bool = False,
max_sequence_length: int = 300,
complex_human_instruction: list[str] | None = None,
)
Encodes the prompt into text encoder hidden states.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
prompt | `str` or `list[str]`, *optional* | prompt to be encoded | required |
negative_prompt | `str` or `list[str]`, *optional* | The prompt not to guide the video generation. If not defined, one has to pass | '' |
do_classifier_free_guidance | `bool`, *optional*, defaults to `True` | whether to use classifier free guidance or not | True |
num_videos_per_prompt | `int`, *optional*, defaults to 1 | number of videos that should be generated per prompt | 1 |
device | device | None | ( | None |
prompt_embeds | `torch.Tensor`, *optional* | Pre-generated text embeddings. Can be used to easily tweak text inputs, e.g. prompt weighting. If not provided, text embeddings will be generated from | None |
negative_prompt_embeds | `torch.Tensor`, *optional* | Pre-generated negative text embeddings. For Sana, it's should be the embeddings of the "" string. | None |
clean_caption | `bool`, defaults to `False` | If | False |
max_sequence_length | `int`, defaults to 300 | Maximum sequence length to use for the prompt. | 300 |
complex_human_instruction | `list[str]`, defaults to `complex_human_instruction` | If | None |
get_sana_video_default_resolution ¶
get_sana_video_post_process_func ¶
get_sana_video_post_process_func(
od_config: OmniDiffusionConfig,
)
resolve_sana_video_sample_size ¶
resolve_sana_video_sample_size(
od_config: OmniDiffusionConfig,
) -> int
retrieve_timesteps ¶
retrieve_timesteps(
scheduler,
num_inference_steps: int | None = None,
device: str | device | None = None,
timesteps: list[int] | None = None,
sigmas: list[float] | None = None,
**kwargs,
) -> tuple[Tensor, int]
Calls the scheduler's set_timesteps method and retrieves timesteps from the scheduler after the call. Handles custom timesteps. Any kwargs will be supplied to scheduler.set_timesteps.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
scheduler | `SchedulerMixin` | The scheduler to get timesteps from. | required |
num_inference_steps | `int` | The number of diffusion steps used when generating samples with a pre-trained model. If used, | None |
device | `str` or `torch.device`, *optional* | The device to which the timesteps should be moved to. If | None |
timesteps | `list[int]`, *optional* | Custom timesteps used to override the timestep spacing strategy of the scheduler. If | None |
sigmas | `list[float]`, *optional* | Custom sigmas used to override the timestep spacing strategy of the scheduler. If | None |
Returns:
| Type | Description |
|---|---|
Tensor |
|
int | second element is the number of inference steps. |