Skip to content

vllm_omni.diffusion.diffusion_kv.request

DiffusionKVContext dataclass

An independently managed K/V context outside the primary sequence.

context_id identifies the logical context within one execution sequence. cache_role binds it to a logical attention cache role exposed by the Worker. Physical cache geometry remains native KVCacheSpec / KVCacheConfig state and is deliberately absent here.

block_hashes class-attribute instance-attribute

block_hashes: tuple[BlockHash, ...] = ()

cache_role instance-attribute

cache_role: str

context_id instance-attribute

context_id: str

num_tokens instance-attribute

num_tokens: int

DiffusionKVRequest

Scheduler-owned KV state for one diffusion execution sequence.

The primary sequence follows an ordered [prefix | target] policy, such as one Hunyuan CFG row. prefix_len is the contiguous reusable prefix; target_len is overwritten by every denoise step; num_tokens is the complete first-step allocation boundary.

Independent cross/joint-attention K/V does not belong to that token axis and is described by kv_contexts. This object also exposes the minimal mutable Request surface consumed by native KVCacheManager.

An empty block_hashes sequence means the prefix has no canonical cache identity yet. Such a request may use native request-local page allocation, but consumers must not publish it through KVCacheManager.cache_blocks. With prefix caching enabled, Engine invokes model input preparation to attach token IDs and native multimodal identities / positions. The Manager builds one canonical hash for every cacheable full block before lookup; publication requires those hashes and successfully materialized KV.

block_hashes instance-attribute

block_hashes = list(block_hashes)

cache_namespace instance-attribute

cache_namespace = cache_namespace

cache_salt instance-attribute

cache_salt = None

cache_token_ids instance-attribute

cache_token_ids = token_ids

kv_contexts instance-attribute

kv_contexts = contexts

kv_transfer_params instance-attribute

kv_transfer_params = kv_transfer_params

lora_request instance-attribute

lora_request = None

mm_features instance-attribute

mm_features = list(mm_features)

num_computed_tokens instance-attribute

num_computed_tokens = 0

num_in_flight_tokens instance-attribute

num_in_flight_tokens = 0

num_preemptions instance-attribute

num_preemptions = 0

num_prompt_tokens instance-attribute

num_prompt_tokens = prefix_len

num_tokens instance-attribute

num_tokens = seq_len

prefix_len instance-attribute

prefix_len = prefix_len

prompt_embeds instance-attribute

prompt_embeds = None

prompt_token_ids instance-attribute

prompt_token_ids = prompt_token_ids

request_id instance-attribute

request_id = request_id

seq_len property

seq_len: int

Complete first-step sequence length and allocation boundary.

sequence_id instance-attribute

sequence_id = sequence_id

shared_prefix_boundary instance-attribute

shared_prefix_boundary = 0

skip_reading_prefix_cache instance-attribute

skip_reading_prefix_cache = not self.block_hashes

status instance-attribute

status = RequestStatus.WAITING

target_len instance-attribute

target_len = target_len

build_block_hashes

build_block_hashes(
    hash_block_size: int,
    hash_function: Callable[[object], bytes],
) -> None

Build native chained hashes for complete reusable-prefix blocks.