vllm_omni.diffusion.diffusion_kv.request ¶
DiffusionKVContext dataclass ¶
An independently managed K/V context outside the primary sequence.
context_id identifies the logical context within one execution sequence. cache_role binds it to a logical attention cache role exposed by the Worker. Physical cache geometry remains native KVCacheSpec / KVCacheConfig state and is deliberately absent here.
DiffusionKVRequest ¶
Scheduler-owned KV state for one diffusion execution sequence.
The primary sequence follows an ordered [prefix | target] policy, such as one Hunyuan CFG row. prefix_len is the contiguous reusable prefix; target_len is overwritten by every denoise step; num_tokens is the complete first-step allocation boundary.
Independent cross/joint-attention K/V does not belong to that token axis and is described by kv_contexts. This object also exposes the minimal mutable Request surface consumed by native KVCacheManager.
An empty block_hashes sequence means the prefix has no canonical cache identity yet. Such a request may use native request-local page allocation, but consumers must not publish it through KVCacheManager.cache_blocks. With prefix caching enabled, Engine invokes model input preparation to attach token IDs and native multimodal identities / positions. The Manager builds one canonical hash for every cacheable full block before lookup; publication requires those hashes and successfully materialized KV.