public R&D · collaboration page · updated 2026-09-10

Scaling Prefix Integrity Analysis

Prefix Integrity Analysis, or PIA, asks whether a saved transformer key/value cache can be reused safely. This work takes that check from an offline screen toward a live serving contract that can survive process boundaries, multiple graphics processors, restarts, and changing model artifacts without confusing one cache object for another.

Current status: useful proof, unfinished system. A bounded one-graphics-processor test stored and restored cache tensors under an externally derived address. It did not run a complete model server or test multiple workers. A commit-by-commit review found correctness defects listed below, so these public branches are for collaboration, not deployment or an upstream merge request.

What “scaling” means here

The original PIA harness cheaply screens one cache transformation at a time. Scaling it has three separate meanings. Keeping them separate prevents a broad experiment from being mistaken for a production-safe serving path.

breadth

More cases

Run the same integrity questions across models, adapters, quantization formats, memory layouts, context lengths, and cache algorithms instead of treating one successful fixture as universal.

depth

Real enforcement

Move from an offline report to the actual model loader, request object, connector, and cache server. Each layer must carry one checked identity without rebuilding a weaker address from tokens.

width

More workers

Make the contract unambiguous when a model is divided across several workers. A worker here means one serving process responsible for one graphics processor or one part of the model.

The failure we want to make impossible

Imagine two requests with identical prompt tokens. One uses the base model and the other uses a fine-tuning adapter, a small set of weights that changes the model. Their tokens match, but their saved attention state may not. If a connector builds an external-cache address from tokens alone, the second request can retrieve the first request's data. The answer is silently wrong and the cache may cross a user boundary.

Our proposed rule is simple: the component that knows what was actually loaded creates the address once. Every later component treats it as an opaque value—meaning it may carry and compare the address, but may not invent one from a smaller set of facts.

1 · resolve

Model loader

Hashes the files actually opened and combines them with request meaning, byte layout, and access scope.

2 · carry

vLLM

Binds the checked addresses to the internal request and passes the exact chunk range to the cache connector.

3 · store

LMCache

Uses the supplied addresses for lookup, storage, retrieval, and lock release without falling back to token hashing.

Public source map

The work spans three repositories because each repository owns a different enforcement point. Branch names are intentionally shown so a collaborator can review the exact work instead of guessing from a default branch.

Repository and branchRolePresent state
modular-kv-pia / main Common cache-block description and stable encoding. Shared base for both PIA and the separate compression experiment.
modular-kv-pia / kv-compat First complete-identity prototype and collision reproductions. Historical parent of the active branch.
modular-kv-pia / resolved contract Registers loaded artifacts, binds requests, and issues opaque cache addresses. Active research prototype
vLLM / provenance contract
5726847ea5
Carries externally issued chunk addresses through the request and multiprocess connector. Known partial-prefix defect
LMCache / provenance contract
4a93ed7b
Consumes those addresses in its multiprocess lookup, storage, retrieval, and locking paths. One-worker path demonstrated
modular-kv-pia / kv-asym-quant Experimental 16-bit-key, 8-bit-value compression codec. Related representation work; not part of the serving-contract stack.

The full Modular branch ancestry and a plain-language description of each branch are in BRANCHES.md.

What the current experiment actually proved

The September 2026 experiment used one NVIDIA L40S graphics processor, PyTorch 2.10.0 with CUDA 12.9, a real LMCache multiprocess server and worker adapter, and selected vLLM request and connector code. The data was synthetic: two layers of 16-bit floating-point cache tensors, one 256-token chunk made of sixteen 16-token vLLM blocks. It was not a full model generation or a vllm serve run.

demonstrated
  • The externally issued address survived serialization between scheduler and worker.
  • LMCache stored one chunk under that address and retrieved it into different graphics-memory blocks.
  • Every restored tensor element exactly matched the stored tensor.
  • The same tokens under a different layout label produced a miss.
  • A token-only legacy lookup did not find the externally addressed object.
  • Repeated lookup and a new scheduler connection found the same object.
not demonstrated
  • No real model loader created and bound the request description.
  • No complete model server generated output from the restored cache.
  • The layout control changed only the label; it did not create and decode two real layouts.
  • No test divided a model across two or more workers.
  • No cross-machine cache transfer, rolling upgrade, or server failure was tested.
  • No latency, throughput, memory-use, or scaling result was measured.

The narrow result is still useful: it proves that the proposed address can cross the two existing multiprocess interfaces and select the intended stored bytes. It is an integration foothold, not evidence that the complete serving design is correct.

What the review found

A commit-by-commit review was run before publishing the integration branches. These are the important defects, stated here so a collaborator does not mistake public source for finished source.

Partial prefix hits can fail in vLLM

If vLLM already has fewer blocks than one complete LMCache chunk, the lock-release path asks for an unaligned external-key range and raises an error. The test covers an aligned hit and misses this ordinary boundary case.

The “immutable” description is only shallowly frozen

The Modular envelope contains mutable dictionaries and lists. A caller can change the layout after the address is calculated, leaving one address attached to different serialized metadata. LMCache has a related problem: its frozen key object accepts a mutable list.

Multi-worker addressing has no consistent contract yet

Modular includes the active worker in its address, vLLM carries one flat address list for the request, and LMCache adds a worker identity separately. The interfaces need one shared rule before a model can be split safely.

Encoding and required-field checks are incomplete

The prototype's Concise Binary Object Representation encoder disagrees with the reference canonical ordering for some supported dictionary keys. The final request-description digest also defaults to an empty value instead of being mandatory.

The production binding does not exist

vLLM exposes a method for binding the addresses, but only tests call it. The real model-loading and request-building path still has to own registration and binding.

The central scaling decision: what does one address name?

When one model is split across several workers, each worker owns different cache bytes. We need to decide whether the externally visible address names the whole logical chunk or one worker's part of that chunk. Both choices can work, but mixing them cannot.

ChoiceHow it worksBenefitCost
One global chunk address Every worker receives the same chunk address. LMCache combines it with a separate worker identifier. Fits vLLM's current flat list and LMCache's current storage key. The external address is not complete by itself; correctness also depends on the hidden worker suffix.
One address per worker part The address includes the exact layers and attention heads owned by one worker. Each address completely identifies the bytes it names. vLLM must carry a worker-to-address mapping rather than one flat list.
Root plus worker child addresses A request-wide root identifies the model, request meaning, access scope, and full partition plan. Deterministic child addresses identify each worker's bytes. Keeps a global audit handle while making every stored object self-contained. Adds a typed mapping and one derivation step to the connector protocol.

Current working direction: evaluate the root-plus-child design first. It preserves a request-wide identity while removing the present ambiguity about which worker's bytes an address names. This is a research direction, not a settled interface.

Ideas for scaling the system

Make descriptions truly immutable

Replace nested dictionaries with typed frozen values, copy data at the boundary, reject empty required fields, and calculate the address only from the frozen form.

Use one reference encoding

Publish byte-for-byte fixtures for Python, Mojo, Rust, and C++. Restrict the supported data types and compare every implementation with an established canonical encoder.

Bind at model load, not at the connector

The loader should register the files it opened and pass a capability-like handle into request construction. Connectors should never accept an operator-written model fingerprint.

Cache descriptions, not trust

Hash artifacts once per load and reuse the resolved registration. Requests derive cheap child addresses from that registered state without rereading model files.

Plan for upgrades and invalidation

Version the wire format and define what happens during a rolling upgrade, an adapter reload, a changed partition plan, or a cache-server restart. Old and new meanings must never share accidentally.

Expose explanations without exposing secrets

Operators need to know why a lookup missed—model changed, layout changed, access scope changed—without logging raw tenant identifiers, prompts, or private artifact paths.

Work list

Correctness comes before scale measurements. A faster ambiguous address is only a faster route to the wrong cache object.

priority 0

Repair the contract

  • Fix vLLM's partial-prefix lock release and add cases below, at, and above one LMCache chunk.
  • Deep-freeze the Modular description and LMCache external-key collection.
  • Correct or deliberately narrow the canonical binary encoder.
  • Make the final request-description digest required and non-empty.
  • Choose one multi-worker address model and make all three repositories implement it.
  • Replace tests that inspect private fields with tests through supported interfaces.
priority 1

Prove the real path

  • Connect registration to the real model and adapter loader, then bind ordinary internal requests automatically.
  • Run a complete vLLM generation that stores, clears, restores, and consumes the cache.
  • Create two genuinely different memory layouts and prove each is decoded correctly, not merely addressed differently.
  • Run a two-worker model split and verify each worker retrieves only its own bytes.
  • Test changed model files, changed adapter files, reloads under the same name, and access-scope separation.
priority 2

Measure scaling

  • Run two, four, and eight workers on one machine, then repeat across machines.
  • Measure address-generation time, message size, lookup latency, throughput, and memory overhead against unmodified vLLM and LMCache.
  • Exercise simultaneous requests, repeated prefixes, partial hits, eviction, reconnects, and rolling process replacement.
  • Expand the PIA matrix across model families, adapter types, cache layouts, and quantization codecs.
priority 3

Prepare an upstream-quality series

  • Rewrite the experimental commit history so every commit is independently correct and its message matches its code.
  • Reduce the public interface to the smallest supported contract and document one concrete end-to-end example.
  • Run each project's complete required checks and record the exact tested hardware and software configuration.
  • Submit separate, reviewable changes only after the cross-project interface is stable.

Good places to collaborate

The most valuable first discussion is the multi-worker address shape. That decision determines the vLLM request interface, the LMCache storage key, and the Modular description at once. Useful review questions are:

  • Should one external address name a whole logical chunk or one worker's part?
  • Which model-loader objects are trustworthy enough to register the loaded artifacts?
  • What is the smallest immutable request description that still covers every cache-changing input?
  • How should old and new wire versions behave during a rolling upgrade?
  • Which failures must be visible to operators without revealing user or model details?

For the identity model, start with the resolved Modular branch. For request transport, read the vLLM branch. For storage and retrieval, read the LMCache branch.