Prefix Integrity Analysis, or PIA, asks whether a saved transformer key/value cache can be reused safely. This work takes that check from an offline screen toward a live serving contract that can survive process boundaries, multiple graphics processors, restarts, and changing model artifacts without confusing one cache object for another.
Current status: useful proof, unfinished system. A bounded one-graphics-processor test stored and restored cache tensors under an externally derived address. It did not run a complete model server or test multiple workers. A commit-by-commit review found correctness defects listed below, so these public branches are for collaboration, not deployment or an upstream merge request.
The original PIA harness cheaply screens one cache transformation at a time. Scaling it has three separate meanings. Keeping them separate prevents a broad experiment from being mistaken for a production-safe serving path.
Run the same integrity questions across models, adapters, quantization formats, memory layouts, context lengths, and cache algorithms instead of treating one successful fixture as universal.
Move from an offline report to the actual model loader, request object, connector, and cache server. Each layer must carry one checked identity without rebuilding a weaker address from tokens.
Make the contract unambiguous when a model is divided across several workers. A worker here means one serving process responsible for one graphics processor or one part of the model.
Imagine two requests with identical prompt tokens. One uses the base model and the other uses a fine-tuning adapter, a small set of weights that changes the model. Their tokens match, but their saved attention state may not. If a connector builds an external-cache address from tokens alone, the second request can retrieve the first request's data. The answer is silently wrong and the cache may cross a user boundary.
Our proposed rule is simple: the component that knows what was actually loaded creates the address once. Every later component treats it as an opaque value—meaning it may carry and compare the address, but may not invent one from a smaller set of facts.
Hashes the files actually opened and combines them with request meaning, byte layout, and access scope.
Binds the checked addresses to the internal request and passes the exact chunk range to the cache connector.
Uses the supplied addresses for lookup, storage, retrieval, and lock release without falling back to token hashing.
The work spans three repositories because each repository owns a different enforcement point. Branch names are intentionally shown so a collaborator can review the exact work instead of guessing from a default branch.
| Repository and branch | Role | Present state |
|---|---|---|
| modular-kv-pia / main | Common cache-block description and stable encoding. | Shared base for both PIA and the separate compression experiment. |
| modular-kv-pia / kv-compat | First complete-identity prototype and collision reproductions. | Historical parent of the active branch. |
| modular-kv-pia / resolved contract | Registers loaded artifacts, binds requests, and issues opaque cache addresses. | Active research prototype |
vLLM / provenance contract5726847ea5 |
Carries externally issued chunk addresses through the request and multiprocess connector. | Known partial-prefix defect |
LMCache / provenance contract4a93ed7b |
Consumes those addresses in its multiprocess lookup, storage, retrieval, and locking paths. | One-worker path demonstrated |
| modular-kv-pia / kv-asym-quant | Experimental 16-bit-key, 8-bit-value compression codec. | Related representation work; not part of the serving-contract stack. |
The full Modular branch ancestry and a plain-language description of each branch are in BRANCHES.md.
The September 2026 experiment used one NVIDIA L40S graphics processor,
PyTorch 2.10.0 with CUDA 12.9, a real LMCache multiprocess server and
worker adapter, and selected vLLM request and connector code. The data
was synthetic: two layers of 16-bit floating-point cache tensors, one
256-token chunk made of sixteen 16-token vLLM blocks. It was not a full
model generation or a vllm serve run.
The narrow result is still useful: it proves that the proposed address can cross the two existing multiprocess interfaces and select the intended stored bytes. It is an integration foothold, not evidence that the complete serving design is correct.
A commit-by-commit review was run before publishing the integration branches. These are the important defects, stated here so a collaborator does not mistake public source for finished source.
If vLLM already has fewer blocks than one complete LMCache chunk, the lock-release path asks for an unaligned external-key range and raises an error. The test covers an aligned hit and misses this ordinary boundary case.
The Modular envelope contains mutable dictionaries and lists. A caller can change the layout after the address is calculated, leaving one address attached to different serialized metadata. LMCache has a related problem: its frozen key object accepts a mutable list.
Modular includes the active worker in its address, vLLM carries one flat address list for the request, and LMCache adds a worker identity separately. The interfaces need one shared rule before a model can be split safely.
The prototype's Concise Binary Object Representation encoder disagrees with the reference canonical ordering for some supported dictionary keys. The final request-description digest also defaults to an empty value instead of being mandatory.
vLLM exposes a method for binding the addresses, but only tests call it. The real model-loading and request-building path still has to own registration and binding.
When one model is split across several workers, each worker owns different cache bytes. We need to decide whether the externally visible address names the whole logical chunk or one worker's part of that chunk. Both choices can work, but mixing them cannot.
| Choice | How it works | Benefit | Cost |
|---|---|---|---|
| One global chunk address | Every worker receives the same chunk address. LMCache combines it with a separate worker identifier. | Fits vLLM's current flat list and LMCache's current storage key. | The external address is not complete by itself; correctness also depends on the hidden worker suffix. |
| One address per worker part | The address includes the exact layers and attention heads owned by one worker. | Each address completely identifies the bytes it names. | vLLM must carry a worker-to-address mapping rather than one flat list. |
| Root plus worker child addresses | A request-wide root identifies the model, request meaning, access scope, and full partition plan. Deterministic child addresses identify each worker's bytes. | Keeps a global audit handle while making every stored object self-contained. | Adds a typed mapping and one derivation step to the connector protocol. |
Current working direction: evaluate the root-plus-child design first. It preserves a request-wide identity while removing the present ambiguity about which worker's bytes an address names. This is a research direction, not a settled interface.
Replace nested dictionaries with typed frozen values, copy data at the boundary, reject empty required fields, and calculate the address only from the frozen form.
Publish byte-for-byte fixtures for Python, Mojo, Rust, and C++. Restrict the supported data types and compare every implementation with an established canonical encoder.
The loader should register the files it opened and pass a capability-like handle into request construction. Connectors should never accept an operator-written model fingerprint.
Hash artifacts once per load and reuse the resolved registration. Requests derive cheap child addresses from that registered state without rereading model files.
Version the wire format and define what happens during a rolling upgrade, an adapter reload, a changed partition plan, or a cache-server restart. Old and new meanings must never share accidentally.
Operators need to know why a lookup missed—model changed, layout changed, access scope changed—without logging raw tenant identifiers, prompts, or private artifact paths.
Correctness comes before scale measurements. A faster ambiguous address is only a faster route to the wrong cache object.
The most valuable first discussion is the multi-worker address shape. That decision determines the vLLM request interface, the LMCache storage key, and the Modular description at once. Useful review questions are:
For the identity model, start with the resolved Modular branch. For request transport, read the vLLM branch. For storage and retrieval, read the LMCache branch.