vllm_omni.diffusion.models.hunyuan_image3.request_layout ¶
HunyuanPreparedLayout dataclass ¶
CPU execution layout prepared once before Scheduler admission.
build_hunyuan_batch_rope_image_info ¶
build_hunyuan_batch_rope_image_info(
output: TokenizerEncodeOutput,
sections: list[list[dict[str, Any]]],
) -> list[list[tuple[slice, tuple[int, int]]]]
build_hunyuan_diffusion_kv_requests ¶
build_hunyuan_diffusion_kv_requests(
request: OmniDiffusionRequest,
prepared_layout: HunyuanPreparedLayout,
) -> tuple[DiffusionKVRequest, ...]
Build allocation-only KV requests, even when prefix caching is disabled.
extract_hunyuan_prompt_inputs ¶
extract_hunyuan_prompt_inputs(
prompts: list[Any],
extra_args: dict[str, Any],
*,
request_id: str,
allow_cond_image: bool,
) -> tuple[
list[str],
list[str | None],
str | None,
list[list[JointImageInfo]] | None,
str,
]
Normalize request prompt fields shared by planning and execution.
get_hunyuan_prepared_layout ¶
get_hunyuan_prepared_layout(
source: Any,
) -> HunyuanPreparedLayout | None
hunyuan_cfg_factor ¶
Return the execution branch count for standard or embedded CFG.
hunyuan_num_image_tokens ¶
Return the generated-image span overwritten on every denoise step.
hunyuan_num_special_tokens ¶
Return the generated-image prefix tokens emitted before latent tokens.
joint_image_info_to_payload ¶
joint_image_info_to_payload(
joint_image_info: JointImageInfo,
) -> dict[str, Any]
native_kv_covers_cond_images ¶
native_kv_covers_cond_images(
output: TokenizerEncodeOutput,
computed_tokens: tuple[int, ...],
) -> bool
Skip image encoding only when every CFG row already contains its image KV.
normalize_hunyuan_cot_text ¶
Restore an AR generation trigger tag omitted from generated text.
normalize_hunyuan_single_stage_bot_task ¶
prepare_hunyuan_layout ¶
prepare_hunyuan_layout(
request: OmniDiffusionRequest,
*,
tokenizer_wrapper: TokenizerWrapper,
image_processor: HunyuanImage3ImageProcessor,
generation_config: GenerationConfig,
image_base_size: int,
cfg_distilled: bool = False,
use_meanflow: bool = False,
) -> HunyuanPreparedLayout
Build the CPU token/image layout reused by Scheduler and Worker.
prepare_hunyuan_prefix_cache ¶
prepare_hunyuan_prefix_cache(
request: OmniDiffusionRequest,
) -> None
Attach native MM identities / positions when Engine enables caching.
This is model input adaptation, not a hashing framework. All reference inputs are hashed once, then reused by VAE/ViT subspans and CFG rows. The native block hasher consumes these ranges directly, without token extras.