Prefix-integrity case study · LMCache + vLLM + SnapKV · September 2026

The data path works.
The cache key is incomplete.

LMCache can run SnapKV against KV produced by a real vLLM engine. Its SDK retrieves KV and query tensors, lets SnapKV select tokens and rerotate keys, stores the compacted result, and sends it back to vLLM for decode. That proves execution—not safe reuse.

A SnapKV tensor depends on the complete source prefix, recent queries, selected original positions, and their new dense positions. Today the multiprocess cache path names model, token chunks, rank, and tenant salt. It does not name that transformation.

Deployment boundaryKeep transformed KV request-local, or isolate it in a separate namespace, until LMCache carries a transform manifest through lookup, storage, and every L2 backend.
REAL VLLM DECODESDK REMAP WORKSMIXED REUSE UNSAFEKEY EXTENSION REQUIRED

What the notebook actually demonstrates

The public LMCache example is an end-to-end inference experiment, not an offline tensor simulation. It intentionally clears the cache between the two arms, so it does not test a plain and a SnapKV-enabled vLLM instance sharing transformed entries.

1 · PrefillvLLM builds the ordinary prompt KV and exposes query tensors through the LMCache connector patch.
2 · RetrieveThe SDK retrieves chunk-aligned KV plus query tensors from the running LMCache service.
3 · Transformdrop_tokens_fn() scores past positions, retains selected tokens, and rerotates keys into dense positions.
4 · DecodeThe SDK stores the shorter KV and vLLM generates from it; accuracy and output throughput are measured.
Execution
YES

Real LMCache shared memory, real vLLM prefill/decode, and the engine's real KV tensors.

Geometry
YES

Logical token count, stored tensor length, retained positions, and RoPE rerotation are updated together.

Mixed sharing
NOT TESTED

Baseline and SnapKV run sequentially with a cache clear between them.

Answer: yes—the notebook explains how to use and test LMCache with vLLM and SnapKV. It validates the token-dropping data path and measures generation. It does not establish that transformed KV is safe to share.

What happens when one vLLM uses SnapKV and another does not

They share entries only if both connect to the same LMCache namespace or shared L2. Separate LMCache services with separate storage do not share. In the shared case, LMCache does not currently encode “ordinary” versus “SnapKV” in this SDK key path.

lookuplikely result todaysafe?reason
Original full prompt after SnapKV modificationOrdinary full-prefix entry, if retainedusuallyThe compacted artifact is stored under its shorter returned token sequence.
Literal prompt equal to retained token IDsMay address transformed entrynoIts ordinary KV is not the KV computed from the longer source and then selected.
Same retained IDs from another source prefixMay address transformed entrynoRetained vectors still encode the omitted source context.
Same source with another queryMay address transformed entrynoSnapKV selection is query-dependent.
Same transform, source, query, parameters and formatPotential reusable hitonly with identityEvery input that defines the payload must match and be verified.
WHAT drop_tokens_fn() FIXES

It gives the SDK a shorter KV tensor and matching logical token list. The example also rerotates retained keys so dense decode positions are internally consistent.

WHAT IT DOES NOT FIX

It does not return a transform identity, source-prefix digest, query digest, retained-position manifest, sharing scope, or compatibility contract.

Temporary isolation: a distinct LMCache service or cache_salt prevents cross-mode hits. Treat that as containment, not the final design: cache_salt represents tenant isolation and quota identity, not artifact compatibility.

Token IDs cannot identify transformed KV

The same logical token sequence can describe different tensors. Prefix integrity requires the cache key to name every input that changes those tensors.

Source history

KV remembers omitted tokens

A retained position was computed while attending over the full original prefix. Removing token IDs from the name does not remove their effect from the vector.

Selection

The query chooses the artifact

SnapKV scores past keys with recent-window queries. Change the query and the selected positions can change while model and source remain the same.

Position

Rerotation changes keys

The example maps retained original positions onto new dense positions. That map and the RoPE contract are part of artifact identity.

Add transform identity, not a special SnapKV exception

Make the API general enough for token dropping, codecs, cartridges, and future KV transforms. Keep tenant isolation separate from compatibility.

Return a structured artifact.Let a modifier return KV, logical token IDs, and a canonical transform manifest. Keep the existing tuple temporarily, but default legacy transformed results to request-local.
Bind every payload-defining input.Include algorithm and version, parameters, complete source-prefix digest, retained original positions, old-to-new position map, query digest when required, model/layout/dtype/RoPE format, and sharing scope.
Derive a transform digest.Carry it through SDK request update, IPCCacheServerKey, ObjectKey, coordinators, events, serializers, directory state, and every L2 adapter.
Require an exact variant at lookup.Ordinary KV uses the empty variant for compatibility. Ordinary lookup never falls back to transformed KV. SnapKV lookup requests the exact manifest digest.
Enforce sharing scope.Use request-local by default; permit query-scoped sharing only when the query digest matches; permit prefix/global sharing only for transforms that prove those narrower inputs fully determine the payload.
producerconsumerrequired result
Ordinary KVOrdinary, same prefixhit
SnapKVOrdinary literal retained-token promptmiss
SnapKV query ASnapKV query A, identical manifesthit
SnapKV query ASnapKV query Bmiss
SnapKV source ASource B with identical retained IDsmiss
SnapKV ratio 0.5SnapKV ratio 0.25miss

Run the negative matrix through process-local L1, multiprocess shared memory, and each enabled L2 backend. Repeat after restart to prove serialized identity survives reload.

Reproduce the real data path from KNLP

KNLP converts the interactive notebook into an unattended, revision-recorded run and exposes it through the kdevops plugin. It measures the current path; it does not label that path safe for mixed reuse.

LOCAL OR PROVISIONED NODE

make defconfig-lmcache-snapkv
make

The harness checks NVIDIA/CUDA prerequisites, fetches LMCache, installs the selected vLLM, applies LMCache's query-tensor patch, and writes a JSON result plus service logs.

KDEVOPS PLUGIN

Enable WORKFLOW_KNLP_LMCACHE_SNAPKV to provision the environment. Enable the separate ..._RUN option to execute the comparison during provisioning.

The default remains build-only so GPU scheduling stays explicit.