Skip to content

vllm_omni.model_extras.registry

ImageToImagePromptBuilder module-attribute

ImageToImagePromptBuilder = Callable[
    [
        str,
        str | None,
        "Image.Image | list[Image.Image]",
        int | None,
        int | None,
    ],
    dict[str, Any],
]

ImageToVideoPromptBuilder module-attribute

ImageToVideoPromptBuilder = Callable[
    [
        str,
        str | None,
        "Mapping[str, Any]",
        int | None,
        int | None,
        int | None,
    ],
    dict[str, Any],
]

OutputTensorRange module-attribute

OutputTensorRange = Literal[
    "negative_one_to_one", "zero_to_one"
]

TextToImagePromptBuilder module-attribute

TextToImagePromptBuilder = Callable[
    [str, str | None, int | None, int | None],
    dict[str, Any],
]

XToTextPromptBuilder module-attribute

XToTextPromptBuilder = Callable[
    [str, str, bool],
    tuple[dict[str, Any], list[int] | None],
]

ReferenceImageSizeResolver

Bases: Protocol

TransformerConfigSubfolderResolver

Bases: Protocol

build_image_to_image_prompt

build_image_to_image_prompt(
    model_class_name: str | None,
    prompt: str,
    negative_prompt: str | None,
    input_image: Image | list[Image],
    height: int | None = None,
    width: int | None = None,
) -> dict[str, Any]

build_image_to_video_prompt

build_image_to_video_prompt(
    model_class_name: str | None,
    prompt: dict[str, Any],
    height: int | None = None,
    width: int | None = None,
    num_frames: int | None = None,
) -> dict[str, Any]

Build a model-specific I2V prompt from an example-owned envelope.

build_text_to_image_prompt

build_text_to_image_prompt(
    model_class_name: str | None,
    prompt: dict[str, Any],
    height: int | None = None,
    width: int | None = None,
) -> dict[str, Any]

Build a model-specific T2I prompt from an example-owned envelope.

build_x_to_text_prompt

build_x_to_text_prompt(
    model_family: str,
    model: str,
    prompt: str,
    has_image: bool,
) -> tuple[dict[str, Any], list[int] | None]

Build a model-aware T2T/I2T prompt and optional stop-token ids.

default_image_to_image_prompt

default_image_to_image_prompt(
    prompt: str,
    negative_prompt: str | None,
    input_image: Image | list[Image],
    height: int | None = None,
    width: int | None = None,
) -> dict[str, Any]

default_x_to_text_prompt

default_x_to_text_prompt(
    model: str, prompt: str, has_image: bool
) -> tuple[dict[str, Any], list[int] | None]

get_ar_input_builder

get_ar_input_builder(
    model_class_name: str | None,
) -> Callable[..., Any] | None

Return a model's AR-stage input builder, or None if undeclared.

Models with a text/AR stage that needs template-formatted prompt tokens and AR stop tokens (e.g. HunyuanImage3) declare an ar_input_builder. The shared task examples call it generically when present, so the example scripts stay model-agnostic; models without one are unaffected.

get_ar_tokenizer_validator

get_ar_tokenizer_validator(
    model_class_name: str | None,
) -> Callable[[Any], None] | None

Return a model's AR-tokenizer validator, or None if undeclared.

Models whose AR prompt/stop-token logic depends on hardcoded special token ids (e.g. HunyuanImage3) declare an ar_tokenizer_validator to check those ids against whatever tokenizer actually loads at runtime. The shared task examples call it generically, right after loading a real tokenizer, so model/tokenizer revision drift fails loudly instead of silently producing a request with the wrong stop tokens.

get_extra_body_params

get_extra_body_params(
    model_class_name: str | None,
) -> frozenset[str]

get_extra_output_params

get_extra_output_params(
    model_class_name: str | None,
) -> frozenset[str]

get_model_class_name

get_model_class_name(omni: Any) -> str | None

Extract model_class_name from an Omni/AsyncOmni instance.

This hides the internal ODConfig plumbing from example scripts.

get_output_tensor_range

get_output_tensor_range(
    model_class_name: str | None,
) -> OutputTensorRange

Return the declared range for floating-point tensor outputs.

The default preserves the shared examples' historical handling. Pipelines that already return normalized tensors declare zero_to_one explicitly.

get_transformer_config_subfolder

get_transformer_config_subfolder(
    model_class_name: str | None,
    *,
    model: str | None,
    revision: str | None = None,
) -> str

Return the model-declared DiT config subfolder, or the standard default.

get_video_generation_defaults

get_video_generation_defaults(
    model_class_name: str | None,
    extra_body: Mapping[str, Any] | None = None,
) -> VideoGenerationDefaults | None

Return model-owned defaults for the shared video examples, if declared.

get_x_to_text_model_family

get_x_to_text_model_family(model: str) -> str

Resolve a text-output prompt family from a checkpoint's config.json.

should_init_extra_args_for_non_diffusion_stages

should_init_extra_args_for_non_diffusion_stages(
    model_class_name: str | None,
) -> bool

should_preserve_reference_image_size

should_preserve_reference_image_size(
    model_class_name: str | None,
    *,
    model: str | None,
    revision: str | None = None,
) -> bool

Return whether the selected pipeline owns reference-image resizing.