Skip to content

vllm_omni.diffusion.diffusion_kv.model_runner_backend

DiffusionKVIdentity module-attribute

DiffusionKVIdentity = tuple[str, int | None, str | None]

DiffusionKVSnapshot module-attribute

DiffusionKVSnapshot = tuple[object, ...]

DiffusionKVModelRunnerBackend

Native paged-KV state and BlockTable operations for a model runner.

block_tables instance-attribute

block_tables: BlockTables | None = None

device instance-attribute

device = device

kv_cache_config instance-attribute

kv_cache_config: KVCacheConfig | None = None

kv_caches instance-attribute

kv_caches: list[Tensor | list[Tensor]] = []

kv_caches_by_layer property

kv_caches_by_layer: dict[str, Tensor]

od_config instance-attribute

od_config = od_config

paged_attention_adapter instance-attribute

paged_attention_adapter: (
    DiffusionPagedAttentionAdapter | None
) = None

vllm_config instance-attribute

vllm_config = vllm_config

activate_paged_attention_metadata

activate_paged_attention_metadata(
    metadata: DiffusionPagedAttentionMetadata,
)

Activate Runner-owned request metadata for an internal denoise loop.

get_diffusion_kv_row

get_diffusion_kv_row(
    request_id: str,
    sequence_id: int | None,
    context_id: str | None = None,
) -> int

Resolve a Scheduler allocation identity to its native table row.

get_paged_attention_adapter

get_paged_attention_adapter() -> (
    DiffusionPagedAttentionAdapter
)

initialize_kv_cache

initialize_kv_cache(kv_cache_config: KVCacheConfig) -> None

Allocate and bind native attention KV tensors for this rank.

install_diffusion_kv_metadata

install_diffusion_kv_metadata(
    metadata: DiffusionKVMetadata,
) -> bool

Install one Scheduler allocation snapshot into native Worker rows.

refresh_block_table_layout

refresh_block_table_layout() -> None

Refresh native pointer tensors after a CuMem KV-cache wake-up.

register_kv_cache_layers

register_kv_cache_layers(
    layers: Mapping[str, tuple[Any, KVCacheSpec]],
) -> dict[str, KVCacheSpec]

Register native adapters and return their canonical cache specs.

remove_diffusion_kv_requests

remove_diffusion_kv_requests(
    request_ids: Sequence[str | tuple[str, int]],
) -> int

Retire Worker rows without logically freeing Scheduler-owned blocks.