vllm_omni.diffusion.diffusion_kv.layout ¶
Physical KV cache layout selection for the native Diffusion paged cache.
vLLM 0.29 removed AttentionSpec.indexes_kv_by_block_stride and replaced it with a single physical layout per model, resolved once and recorded on CacheConfig.kv_cache_layout. KVCacheLayout.is_block_outermost is the direct successor of the old per-spec flag: upstream derives interleaved_block_stride from it when it lays out KVCacheTensor regions (vllm/v1/core/kv_cache_utils.py).
Diffusion never reaches vLLM's engine-core resolution. Its backends subclass Omni's own AttentionBackend ABC, which upstream's get_supported_kv_cache_layouts never sees, so the layout is pinned here before get_kv_cache_configs runs.
BLOCK_STRIDE_DIFFUSION_KV_CACHE_LAYOUT module-attribute ¶
DEFAULT_DIFFUSION_KV_CACHE_LAYOUT module-attribute ¶
adopt_kv_cache_layout ¶
Adopt the control plane's resolved layout inside a worker.
The engine core resolves the layout once and every worker adopts it before allocating (CacheConfig.kv_cache_layout). Diffusion workers receive only a KVCacheConfig, so the name rides along on kv_cache_config and is copied onto the worker's own config here. Without this, anything calling get_resolved_kv_cache_layout in a worker -- init_kv_cache does -- raises.
assert_backend_layout_supported ¶
assert_backend_layout_supported(
vllm_config: VllmConfig, attn_backend: type | None
) -> None
Fail loudly when a backend's block-stride need contradicts the layout.
Skipped while the layout is still unresolved -- worker processes collect specs before the control plane hands them the resolved name -- so this enforces the contract wherever it is knowable and never crashes a rank that has simply not been told yet.
build_kv_cache_tensor ¶
build_kv_cache_tensor(
spec: KVCacheSpec,
num_blocks: int,
layers: Sequence[str],
*,
layout: KVCacheLayout | None = None,
offset: int = 0,
) -> KVCacheTensor
A KVCacheTensor covering layers for num_blocks blocks.
Before 0.29 a tensor named the layers sharing it (KVCacheTensor(size=..., shared_by=[...])); 0.29 replaced that with an explicit layout, so layer l's block b starts at offset + l * layer_stride + b * block_stride. Note this lays the layers out at distinct offsets rather than aliasing them, which is what upstream now does for a single cache group.
The strides come from upstream's own compute_layout_strides, called the way vllm.v1.core.kv_cache_utils calls it, because hand-computed strides read K/V from the wrong offsets silently instead of raising.
get_connector_required_kv_cache_layout ¶
Resolve the physical layout declared by the configured native connector.
resolve_diffusion_kv_cache_layout ¶
resolve_diffusion_kv_cache_layout(
vllm_config: VllmConfig,
*,
indexes_kv_by_block_stride: bool = False,
required_layout: KVCacheLayout | str | None = None,
) -> KVCacheLayout
Pin the physical KV layout for the Diffusion cache and return it.
A layout already present on the config wins, mirroring upstream's resolve_kv_cache_layout; it is validated against the requirement rather than silently overwritten, because an inconsistent layout reads K/V from the wrong offsets instead of raising.