Skip to content

vllm_omni.diffusion.models.minimax_h3.quality_policy

Model-owned request quality policy for MiniMax H3.

MINIMAX_H3_FORCE_REFRESH_POLICIES module-attribute

MINIMAX_H3_FORCE_REFRESH_POLICIES = ('once', 'repeat')

MINIMAX_H3_FORCE_REFRESH_STEP_HINT_ARG module-attribute

MINIMAX_H3_FORCE_REFRESH_STEP_HINT_ARG = (
    "force_refresh_step_hint"
)

MINIMAX_H3_FORCE_REFRESH_STEP_POLICY_ARG module-attribute

MINIMAX_H3_FORCE_REFRESH_STEP_POLICY_ARG = (
    "force_refresh_step_policy"
)

MINIMAX_H3_GENERIC_CACHE_KEY module-attribute

MINIMAX_H3_GENERIC_CACHE_KEY = 'minimax_h3.generic'

MINIMAX_H3_HIGH_CACHE_KEY module-attribute

MINIMAX_H3_HIGH_CACHE_KEY = 'minimax_h3.high'

MiniMaxH3QualityPlan dataclass

Resolved execution choices for one MiniMax H3 request.

cache_dit instance-attribute

cache_dit: CacheDiTRequestSpec | None

MiniMaxH3QualityPolicy

Resolve H3 request quality into a declarative Cache-DiT target.

When the server starts with Cache-DiT, omitted quality selects the server-configured profile. Otherwise, omitted quality selects no cache. In either case, lossless selects no cache and high selects H3's high-quality profile. The policy therefore owns whether a request installs Cache-DiT; the startup backend only controls the omitted-quality default. H3-specific Cache-DiT refresh hints are read from request extra_args; they do not change the global diffusion request contract. The pipeline owns applying the resulting target at the request boundary.

resolve

resolve(
    *,
    quality: str | None,
    num_inference_steps: int,
    extra_args: Mapping[str, Any] | None = None,
) -> MiniMaxH3QualityPlan