Skip to content

vllm_omni.diffusion.models.sana_wm.request

Request normalization for Sana-WM image-to-video.

The released SANA-WM checkpoint is first-frame image-to-video. This module keeps request-side normalization lightweight: it validates the image/camera/action contract and stores a canonical payload under additional_information["sana_wm"]. Model-internal raymap / Plucker projection lives with the transformer.

This lives beside the model rather than under stage_input_processors because nothing loads it as one: the pipeline's pre_process_func is registered from pipeline_sana_wm, the same as every other diffusion model, and no deploy config wires a custom_process_input_func here. Keeping it in the stage package forced the pipeline to import upwards — the only such import in diffusion/models/ — which in turn made the VAE compression constants impossible to share and so duplicated.

SANA_WM_CANONICAL_KEY module-attribute

SANA_WM_CANONICAL_KEY = 'sana_wm'

SANA_WM_DEFAULT_CAMERA_FORMAT module-attribute

SANA_WM_DEFAULT_CAMERA_FORMAT = 'c2w_4x4'

SANA_WM_DEFAULT_COORDINATE_SYSTEM module-attribute

SANA_WM_DEFAULT_COORDINATE_SYSTEM = 'official'

SANA_WM_DEFAULT_HEIGHT module-attribute

SANA_WM_DEFAULT_HEIGHT = 704

SANA_WM_DEFAULT_NUM_FRAMES module-attribute

SANA_WM_DEFAULT_NUM_FRAMES = 161

SANA_WM_DEFAULT_WIDTH module-attribute

SANA_WM_DEFAULT_WIDTH = 1280

normalize_sana_wm_payload

normalize_sana_wm_payload(
    prompt: Mapping[str, Any],
) -> dict[str, Any]

Return a prompt copy with canonical Sana-WM request metadata.

The sana_wm block is read from the top-level sana_wm key, falling back to additional_information["sana_wm"] so the function is idempotent when called again on its own output. The first-frame image is read from multi_modal_data["image"].