Skip to content

vllm_omni.model_extras.hunyuan_image3

HUNYUAN_IMAGE3_EXTRA_BODY_PARAMS module-attribute

HUNYUAN_IMAGE3_EXTRA_BODY_PARAMS = frozenset(
    {
        "bot_task",
        "use_system_prompt",
        "system_prompt",
        "negative_prompt",
    }
)

HUNYUAN_IMAGE3_EXTRA_OUTPUT_PARAMS module-attribute

HUNYUAN_IMAGE3_EXTRA_OUTPUT_PARAMS = frozenset()

HUNYUAN_IMAGE3_INIT_EXTRA_ARGS_FOR_NON_DIFFUSION_STAGES module-attribute

HUNYUAN_IMAGE3_INIT_EXTRA_ARGS_FOR_NON_DIFFUSION_STAGES = (
    True
)

ARStageInputs dataclass

AR-stage inputs the shared example applies to a HunyuanImage3 request.

prompt_token_ids (byte-for-byte HF parity) is preferred; prompt is the string fallback used when no tokenizer is available. modalities picks the AR output type (["image"] for generation, ["text"] for comprehension). use_system_prompt feeds the DiT system-prompt prefix and stop_token_ids terminate the AR decode.

stage_indices names which stage(s) in the request's sampling_params_list these stop_token_ids belong to -- the shared scripts apply them by explicit index, not by scanning for "whichever stage isn't a diffusion stage." HunyuanImage-3.0's topology always puts its single AR stage at index 0 (see vllm_omni.model_executor.models.hunyuan_image3.pipeline); a future model with multiple AR/understanding stages would declare its own indices here instead of relying on type-based inference.

modalities instance-attribute

modalities: list[str]

prompt instance-attribute

prompt: str | None

prompt_token_ids instance-attribute

prompt_token_ids: list[int] | None

stage_indices class-attribute instance-attribute

stage_indices: list[int] = field(
    default_factory=lambda: [0]
)

stop_token_ids instance-attribute

stop_token_ids: list[int]

use_system_prompt instance-attribute

use_system_prompt: str

build_ar_stage_inputs

build_ar_stage_inputs(
    prompt: str,
    tokenizer: Any | None,
    extra_body: Mapping[str, Any] | None,
    *,
    num_images: int = 0,
    height: int | None = None,
    width: int | None = None,
    text_output: bool = False,
) -> ARStageInputs

Resolve HunyuanImage-3.0 AR-stage inputs declaratively.

Invoked generically by the shared task examples (text_to_image / image_to_image / x_to_text) whenever the model declares an ar_input_builder. All model-specific knobs (bot_task / use_system_prompt / system_prompt) arrive via extra_body, so the example itself stays model-agnostic. The heavy lifting is delegated to :func:build_ar_prompt_inputs. The OpenAI server's serving_chat.py does not go through this seam -- it independently calls the same underlying build_prompt/build_prompt_tokens/resolve_stop_token_ids primitives (see that function's docstring for the resulting divergence risk).

build_x_to_text_prompt

build_x_to_text_prompt(
    model: str, prompt: str, has_image: bool
) -> tuple[dict[str, Any], list[int] | None]

Build HunyuanImage-3 prompt tokens and stopping rules for T2T/I2T.

Routes through :func:build_ar_stage_inputs, the same declarative seam the shared text_to_image / image_edit examples use, instead of re-deriving the AR prefill and stop-token logic here. Unlike those examples -- generic scripts that reach validate_ar_tokenizer only indirectly via get_ar_tokenizer_validator -- this function is already HunyuanImage3-specific, so it calls the validator directly rather than requiring the (model-agnostic) x_to_text.py caller to know about it.

validate_ar_tokenizer

validate_ar_tokenizer(tokenizer: Any) -> None

Registry-declared ar_tokenizer_validator hook for HunyuanImage3.

Called by the shared task examples right after they load a real tokenizer for the AR stage, so a model/tokenizer revision bump that silently shifts special-token ids fails loudly instead of producing a request that completes with the wrong stop tokens.