The data path works.
The cache key is incomplete.
LMCache can run SnapKV against KV produced by a real vLLM engine. Its SDK retrieves KV and query tensors, lets SnapKV select tokens and rerotate keys, stores the compacted result, and sends it back to vLLM for decode. That proves execution—not safe reuse.
A SnapKV tensor depends on the complete source prefix, recent queries, selected original positions, and their new dense positions. Today the multiprocess cache path names model, token chunks, rank, and tenant salt. It does not name that transformation.
What the notebook actually demonstrates
The public LMCache example is an end-to-end inference experiment, not an offline tensor simulation. It intentionally clears the cache between the two arms, so it does not test a plain and a SnapKV-enabled vLLM instance sharing transformed entries.
drop_tokens_fn() scores past positions, retains selected tokens, and rerotates keys into dense positions.Real LMCache shared memory, real vLLM prefill/decode, and the engine's real KV tensors.
Logical token count, stored tensor length, retained positions, and RoPE rerotation are updated together.
Baseline and SnapKV run sequentially with a cache clear between them.
Token IDs cannot identify transformed KV
The same logical token sequence can describe different tensors. Prefix integrity requires the cache key to name every input that changes those tensors.
KV remembers omitted tokens
A retained position was computed while attending over the full original prefix. Removing token IDs from the name does not remove their effect from the vector.
The query chooses the artifact
SnapKV scores past keys with recent-window queries. Change the query and the selected positions can change while model and source remain the same.
Rerotation changes keys
The example maps retained original positions onto new dense positions. That map and the RoPE contract are part of artifact identity.
Add transform identity, not a special SnapKV exception
Make the API general enough for token dropping, codecs, cartridges, and future KV transforms. Keep tenant isolation separate from compatibility.
IPCCacheServerKey, ObjectKey, coordinators, events, serializers, directory state, and every L2 adapter.| producer | consumer | required result |
|---|---|---|
| Ordinary KV | Ordinary, same prefix | hit |
| SnapKV | Ordinary literal retained-token prompt | miss |
| SnapKV query A | SnapKV query A, identical manifest | hit |
| SnapKV query A | SnapKV query B | miss |
| SnapKV source A | Source B with identical retained IDs | miss |
| SnapKV ratio 0.5 | SnapKV ratio 0.25 | miss |
Run the negative matrix through process-local L1, multiprocess shared memory, and each enabled L2 backend. Repeat after restart to prove serialized identity survives reload.
Reproduce the real data path from KNLP
KNLP converts the interactive notebook into an unattended, revision-recorded run and exposes it through the kdevops plugin. It measures the current path; it does not label that path safe for mixed reuse.
make defconfig-lmcache-snapkvmake
The harness checks NVIDIA/CUDA prerequisites, fetches LMCache, installs the selected vLLM, applies LMCache's query-tensor patch, and writes a JSON result plus service logs.
Enable WORKFLOW_KNLP_LMCACHE_SNAPKV to provision the environment. Enable the separate ..._RUN option to execute the comparison during provisioning.
The default remains build-only so GPU scheduling stays explicit.