Skip to content

vllm_omni.diffusion.diffusion_kv.kv_cache_utils

Small diffusion identity helpers alongside vLLM's native KV cache utilities.

Content hashing runs before Scheduler admission, only with prefix caching on. Requests carry native MultiModalFeatureSpec / PlaceholderRange values; native generate_block_hash_extra_keys handles block intersections, with no per-token extra-key expansion or intermediate dependency objects.

get_cache_namespace

get_cache_namespace(
    model_namespace: str,
    sampling: OmniDiffusionSamplingParams,
) -> str

Isolate model semantics and the adapter actually activated by Worker.

Native LoRA block keys use lora_name, but diffusion also supports a scale and identifies loaded adapters by lora_int_id. Keep that difference here. Seeds are NOT global identity inputs: models attach random state only to the multimodal features whose computation depends on it.

hash_prefix_cache_value

hash_prefix_cache_value(value: object) -> bytes

Deterministic, type-framed identity using vLLM's multimodal hasher.