vllm_omni.core.prefix_cache.interface ¶
Interface types for the omni prefix cache.
Naming aligns with vLLM's v1/core KV-cache design.
ModelCachePolicy dataclass ¶
Replaces getattr probing on models for cache behavior decisions.
Hidden's name is HIDDEN_KEY (shared identity). This object answers whether this model caches it, and which mm keys are deferred.
deferred_keys class-attribute instance-attribute ¶
deferred_keys: frozenset[TensorName] = frozenset()
hidden_key property ¶
hidden_key: TensorName | None
Pool key for hidden, or None when this model opts out.
from_model classmethod ¶
from_model(model: Any) -> ModelCachePolicy
Shim over legacy per-model attributes (deprecation window).
get_hit_keys ¶
get_hit_keys(
keys: Iterable[TensorName],
) -> list[TensorName]
Keys to plan/prefetch for a hit: hidden first (if cached), then mm.
skip_immediate_mm ¶
skip_immediate_mm(key: TensorName) -> bool
Immediate on-device clone must not take hidden or deferred keys from mm.
OmniPrefixCacheStagingTimeoutError ¶
Bases: OmniPrefixCacheUnmatchError
Save waited for a free staging slot and timed out.
OmniPrefixCacheUnmatchError ¶
Bases: RuntimeError
Contract, config, or KV-occupancy error that must raise.
Includes hit spans that resolve to absent slots (omni cache diverged from vLLM KV), a step id consumed twice or never saved, a step larger than the staging page, and a failed write. Do not pretend these were a miss.
PrefixCacheConfig dataclass ¶
Sizing and flow-control knobs (mirrors KVCacheConfig).
from_vllm_config classmethod ¶
from_vllm_config(
*,
num_blocks: int,
block_size: int,
scheduler_config: Any = None,
model_config: Any = None,
) -> PrefixCacheConfig
Size device→host staging from the running scheduler.
A slot holds one step (the whole batch), not one request: staging_capacity_tokens is max_num_batched_tokens (falls back to max_model_len); staging_depth is how many unconsumed step ids may exist at once, not max_num_seqs.
Pinned staging is allocated lazily per key at depth * capacity_tokens * width * dtype. There is no clamp: a step larger than capacity raises. A 16k-token thinking batch at hidden=2048 bf16 is ~256 MiB for hidden alone; each mm key adds another pool tensor.
staging_depth is the dataclass default (4). There is no CLI or deploy YAML knob — changing it is a code change. Every save that issues a step id claims one slot, including saves with only leftover mm. A full pool waits for materialize/discard; staging_claim_timeout_s then errors.
StageCacheOutputs ¶
Bases: NamedTuple
Plain value object: per-request merged stage outputs.
WriteSchedule ¶
without_hidden ¶
without_hidden(
keys: Iterable[TensorName],
) -> set[TensorName]
Drop the reserved hidden identity; the rest are mm names.