vllm_omni.diffusion.diffusion_kv.kv_cache_utils ¶
Small diffusion identity helpers alongside vLLM's native KV cache utilities.
Content hashing runs before Scheduler admission, only with prefix caching on. Requests carry native MultiModalFeatureSpec / PlaceholderRange values; native generate_block_hash_extra_keys handles block intersections, with no per-token extra-key expansion or intermediate dependency objects.
get_cache_namespace ¶
get_cache_namespace(
model_namespace: str,
sampling: OmniDiffusionSamplingParams,
) -> str
Isolate model semantics and the adapter actually activated by Worker.
Native LoRA block keys use lora_name, but diffusion also supports a scale and identifies loaded adapters by lora_int_id. Keep that difference here. Seeds are NOT global identity inputs: models attach random state only to the multimodal features whose computation depends on it.