Skip to content

vllm_omni.core.prefix_cache.interface

Interface types for the omni prefix cache.

Naming aligns with vLLM's v1/core KV-cache design.

HIDDEN_KEY module-attribute

HIDDEN_KEY: TensorName = '__hidden_states__'

ReqId module-attribute

ReqId: TypeAlias = str

StepId module-attribute

StepId: TypeAlias = int

TensorName module-attribute

TensorName: TypeAlias = str

Tid module-attribute

Tid: TypeAlias = int

ModelCachePolicy dataclass

Replaces getattr probing on models for cache behavior decisions.

Hidden's name is HIDDEN_KEY (shared identity). This object answers whether this model caches it, and which mm keys are deferred.

deferred_keys class-attribute instance-attribute

deferred_keys: frozenset[TensorName] = frozenset()

hidden_key property

hidden_key: TensorName | None

Pool key for hidden, or None when this model opts out.

needs_full_hidden_states class-attribute instance-attribute

needs_full_hidden_states: bool = True

from_model classmethod

from_model(model: Any) -> ModelCachePolicy

Shim over legacy per-model attributes (deprecation window).

get_hit_keys

get_hit_keys(
    keys: Iterable[TensorName],
) -> list[TensorName]

Keys to plan/prefetch for a hit: hidden first (if cached), then mm.

skip_immediate_mm

skip_immediate_mm(key: TensorName) -> bool

Immediate on-device clone must not take hidden or deferred keys from mm.

OmniPrefixCacheStagingTimeoutError

Bases: OmniPrefixCacheUnmatchError

Save waited for a free staging slot and timed out.

OmniPrefixCacheUnmatchError

Bases: RuntimeError

Contract, config, or KV-occupancy error that must raise.

Includes hit spans that resolve to absent slots (omni cache diverged from vLLM KV), a step id consumed twice or never saved, a step larger than the staging page, and a failed write. Do not pretend these were a miss.

PrefixCacheConfig dataclass

Sizing and flow-control knobs (mirrors KVCacheConfig).

block_size instance-attribute

block_size: int

copy_chunk_bytes class-attribute instance-attribute

copy_chunk_bytes: int = 16 * 1024 * 1024

gpu_staging_bytes class-attribute instance-attribute

gpu_staging_bytes: int = 512 * 1024 * 1024

num_blocks instance-attribute

num_blocks: int

staging_capacity_tokens class-attribute instance-attribute

staging_capacity_tokens: int = 1024

staging_claim_timeout_s class-attribute instance-attribute

staging_claim_timeout_s: float = 30.0

staging_depth class-attribute instance-attribute

staging_depth: int = 4

from_vllm_config classmethod

from_vllm_config(
    *,
    num_blocks: int,
    block_size: int,
    scheduler_config: Any = None,
    model_config: Any = None,
) -> PrefixCacheConfig

Size device→host staging from the running scheduler.

A slot holds one step (the whole batch), not one request: staging_capacity_tokens is max_num_batched_tokens (falls back to max_model_len); staging_depth is how many unconsumed step ids may exist at once, not max_num_seqs.

Pinned staging is allocated lazily per key at depth * capacity_tokens * width * dtype. There is no clamp: a step larger than capacity raises. A 16k-token thinking batch at hidden=2048 bf16 is ~256 MiB for hidden alone; each mm key adds another pool tensor.

staging_depth is the dataclass default (4). There is no CLI or deploy YAML knob — changing it is a code change. Every save that issues a step id claims one slot, including saves with only leftover mm. A full pool waits for materialize/discard; staging_claim_timeout_s then errors.

StageCacheOutputs

Bases: NamedTuple

Plain value object: per-request merged stage outputs.

hidden_states instance-attribute

hidden_states: dict[ReqId, Tensor] | None

mm_outputs instance-attribute

mm_outputs: dict[TensorName, dict[ReqId, Any]]

WriteSchedule

Bases: Enum

Write scheduling policy for one WriteTask.

JOIN_NEXT_STEP class-attribute instance-attribute

JOIN_NEXT_STEP = 'join_next_step'

JOIN_ON_FINISH class-attribute instance-attribute

JOIN_ON_FINISH = 'join_on_finish'

is_hidden_key

is_hidden_key(key: TensorName) -> bool

without_hidden

without_hidden(
    keys: Iterable[TensorName],
) -> set[TensorName]

Drop the reserved hidden identity; the rest are mm names.