vllm_omni.model_executor.models.model_local_kv ¶
Declaration protocol for attention KV kept outside the paged manager.
Several models hold their own attention KV as HuggingFace transformers cache objects rather than using the engine's paged KV manager. That memory is allocated after the profiling run that decides how much KV the engine may claim, so no per-stage footprint number accounts for it, and nothing catches a cache that silently stops working.
This module lets a model declare what it holds. It describes; it never allocates. Allocation for the diffusion path is owned by the DiT KV manager (RFC #5244 / PR #6094) and this protocol must not grow into a second allocator.
Why a runtime query rather than a static table: geometry is frequently not knowable until weights are loaded. ming_flash_omni's talker builds its cache from a Qwen2Config that comes from the checkpoint, so layers, kv-heads, head-dim and dtype simply do not exist in this repo. A table would have a hole in it; a post-load query does not.
Why capacity is a resolved number rather than a formula: the four known caches are each bounded by a different mechanism -- a sliding window, a decode-loop trip count, an encoder frame limit, a hardcoded constant -- and not one of them is max_model_len. Only the model knows its own bound.
HasModelLocalKV ¶
Bases: Protocol
Implemented by whichever object owns the cache.
Implement it on the owner, not on the registered model. The owner is usually several levels down (a codec decoder, a talker backbone), and collect_model_local_kv_specs finds it by walking the module tree, so no intermediate class has to forward anything.
ModelLocalKVScope ¶
How long one allocation lives.
Diagnostic only. Lifetime and size are independent: a per-call cache can be the widest thing a model owns, and a process-lifetime one can be a single row. ModelLocalKVSpec.rows carries the size; this carries only the reader's mental model of when the memory appears and goes.
INVOCATION class-attribute instance-attribute ¶
Dies when the call returns (e.g. a per-step working copy).
MODEL class-attribute instance-attribute ¶
Lives as long as the model (e.g. captured into a CUDA-graph pool).
REQUEST class-attribute instance-attribute ¶
Retained across steps of one request, released when it finishes.
SESSION class-attribute instance-attribute ¶
Outlives a request; belongs to a duplex/streaming session.
ModelLocalKVSpec dataclass ¶
One declared KV allocation.
A model returns one entry per distinct lifetime, not per cache object: if the same geometry exists as both a retained per-request cache and a short-lived working copy, that is two entries.
allocation_note class-attribute instance-attribute ¶
allocation_note: str | None = None
Optional detail such as "rebuilt per text segment". Diagnostic only.
capacity_source instance-attribute ¶
capacity_source: str
Free-text note on where the bound comes from, for diagnostics only.
Never branch on this. It exists so a reader can tell "71 because sliding window" from "2048 because someone typed 2048", which is otherwise invisible.
physical_capacity_positions instance-attribute ¶
physical_capacity_positions: int
Positions the tensor can physically hold.
Not the logical sequence position. A sliding-window cache truncates on every write, so a stream of thousands of tokens may only ever occupy sliding_window - 1 slots.
rows instance-attribute ¶
rows: RowDriver
What sets how many rows of this shape are live at peak.
A row is one batch entry's worth of the geometry above. Deliberately does not distinguish "one allocation N rows wide" from "N allocations of one row": those cost the same, and an earlier revision that tried to carry both got the object topology wrong in three of four declarers. Bytes are what this protocol reports; describe the layout in allocation_note.
Required, because defaulting it silently understates every cache whose multiplicity is not 1.
rows_fixed class-attribute instance-attribute ¶
rows_fixed: int = 1
The count, when rows is FIXED. Ignored otherwise.
rows_reason class-attribute instance-attribute ¶
rows_reason: str = ''
Why rows_fixed is that number. Required when rows is FIXED.
A bare constant cannot be reviewed. rows_fixed=1 on a cache that looks per-request is exactly the claim a reader has to be able to check.
scope instance-attribute ¶
scope: ModelLocalKVScope
How long one allocation lives. Diagnostic only.
Deliberately not the multiplier. An earlier revision derived scaling from scope, which made every non-MODEL cache multiply by max_num_seqs. That over-reported ming by that factor (its calls are serialized) and used the wrong driver entirely for MiniCPM-o (duplex sessions, not sequences). Lifetime does not imply replication topology; rows says what does.
peak_bytes ¶
Peak bytes once the engine's concurrency is known.
The model supplies geometry and names the driver; only the engine knows the number behind it.
RowDriver ¶
What sets how many rows of a cache are live at peak.
Two values, because two are all the known caches need. A third belongs here only when a model actually has a cache the engine widens by some other number -- adding one speculatively costs a branch everywhere and buys nothing.
collect_model_local_kv_specs ¶
collect_model_local_kv_specs(
model: object,
) -> list[tuple[str, ModelLocalKVSpec]]
Collect every declaration in a loaded model, with the owner's path.
Walks named_modules() because the cache owner is an inner module in all known cases. Requiring each registered model to forward the call would put the burden on classes that do not own a cache and would silently report zero the moment someone forgot -- exactly the failure mode this is meant to surface.
A raising declaration is logged and skipped rather than propagated: reporting memory must not be able to break model load.
spec_from_hf_config ¶
spec_from_hf_config(
config: object,
*,
name: str,
physical_capacity_positions: int,
capacity_source: str,
scope: ModelLocalKVScope,
dtype: dtype,
rows: RowDriver,
rows_fixed: int = 1,
rows_reason: str = "",
allocation_note: str | None = None,
layers: int | None = None,
kv_heads: int | None = None,
head_dim: int | None = None,
) -> ModelLocalKVSpec
Build a spec from the same config object the cache is built from.
Every known model-local cache is constructed from an HF config, so the geometry half of a declaration is the same three lookups each time. Pass config and this derives them; pass layers/kv_heads/head_dim explicitly to override when the config uses encoder-decoder naming or when the built module is a more truthful source than the config.
Raises rather than guessing when an attribute is absent: a silently wrong geometry would under-report, which is the failure this protocol exists to prevent. rows is keyword-required for the same reason -- an earlier revision defaulted the batch extent to 1 here while documenting it as required on the dataclass, so every declarer that used this helper got the default and the guarantee was decorative.