Cartridge deployment · ecosystem roadmap · September 2026

Route the artifact.
Keep the engine in control.

Cartridge support can begin with an orchestrated cache transform and grow into direct loading of trained artifacts. A Cartridge is a learned KV prefix that carries document knowledge. LMCache can prefill an ordinary prompt, pass its real KV through drop_tokens_fn(), and replay the transformed cache through vLLM. A complementary path loads an offline-trained Cartridge by stable artifact ID and skips document prefill.

Both paths can share routing, storage, residency, and inference infrastructure. The transformation path is a practical first ecosystem integration; direct loading adds artifact identity, virtual positions, and ordered composition for Cartridges at Scale (CAS).

Recommended sequenceStart with LMCache's drop_tokens_fn() orchestration for prompt-derived Cartridge experiments. Extend the same data plane with opaque artifact IDs, router hints, virtual-prefix allocation, and a portable manifest for direct trained-artifact loading and CAS composition.
TRAIN ONCEROUTE BY ARTIFACT IDK16/V8 READYCAS COMPOSITION OPEN

Two complementary paths into the same serving ecosystem

Both paths move KV through a real inference engine. They differ in where the object comes from, whether document prefill is avoided, and what must be present in the cache key.

Orchestrated cache transform

Prefill, transform, replay

  1. vLLM prefills the original prompt.
  2. LMCache retrieves its KV and token IDs.
  3. drop_tokens_fn() returns edited KV and a matching token sequence.
  4. LMCache stores the result and vLLM decodes from it.

This is the LMCache path proposed as the first outlet for Cartridge implementations. The transform may be query-aware, as in SnapKV, or query-independent. It can change the number of positions and rerotate retained keys into a dense sequence.

Cost shape. This route reduces later decode memory, but it does not eliminate the first prefill that produced the cache.
Direct trained-artifact load

Select, fetch, attach

  1. A router resolves the request to one Cartridge ID or an ordered set.
  2. A registry validates the immutable artifact manifest.
  3. The data plane makes the payload resident.
  4. vLLM allocates virtual prefix positions and the attention backend reads them.

The source document never needs to appear in the request. The artifact has its own identity, length, position map, model compatibility, and K/V representation.

Cost shape. Reuse skips source-document prefill and amortizes offline training over every request that loads the Cartridge.
The cache key is not just the returned token list. A transformed KV vector still encodes its original prefix. Query-aware selection also encodes the query. Safe reuse must bind the source-prefix digest, query dependency, transform version and parameters, retained positions, RoPE mapping, model layout, and payload digest.

Give each layer one job

The router chooses meaning. The data plane moves bytes. vLLM owns execution state. FlashInfer turns the representation into one attention result.

RouterMap the request to an ordered Cartridge set and placement hints.
RegistryResolve immutable IDs to manifests, versions, digests, and locations.
Data planeFetch, prefetch, tier, pin, verify, and report residency.
vLLMAllocate exact virtual positions, schedule the request, and offset live tokens.
AttentionRead Cartridge plus live KV with the correct dtype and one softmax state.
ROUTER SELECTS IDs

The request can name an artifact explicitly or let a semantic router select it. File paths and GPU pages stay below this interface.

ENGINE OWNS MEMORY

The connector fills blocks allocated by vLLM. It does not compete with the scheduler for GPU pages or invent a second allocation model.

KERNEL OWNS MIXED DTYPE

K16/V8 is an artifact representation. FlashInfer loads BF16 keys and FP8 values while the request's live cache may remain BF16.

The components exist; the complete ecosystem does not yet

Base Cartridges compile one document into one learned KV prefix. Cartridges at Scale (CAS) extends the design toward collections and jointly compatible composition; the current reproduction validates isolated training, not co-loaded serving. Asymmetric Cartridge KV is a post-training representation that keeps keys in BF16 and stores values in FP8. The table separates those measured pieces from the joins that still need end-to-end proof.

componenttested resultcurrent limit
Base CartridgesTraining, quality, direct KV injection, per-request dispatch, isolation, residency, eviction, and tensor-parallel shardingExisting connector replaces prompt positions; a portable virtual-prefix contract remains to be standardized.
CAS reproductionFaithful isolated training, five-document evaluation, full-document cache-path control, and training-stability measurementsOrdered multi-Cartridge composition and joint mixed-visibility serving are not yet a production result.
Asymmetric Cartridge KVK16/V8 quality, physical FP8 value storage, exact-length prefill and decode, and fused H100 serving against BF16 Cartridge injectionOne homogeneous Cartridge per request group; broader GPUs, workloads, graph execution, and multi-artifact fusion remain open.
LMCache transformationVariable-length edited KV, matching token IDs, query transfer, SnapKV/R-KV examples, and real-engine replayExamples use offline batched orchestration and begin from an ordinary prefill.
LMCache Cartridge storageRead-only storage-plugin integration and connector-side testsThe storage plugin is not a complete standard for trained-artifact routing and virtual-prefix semantics.
KVCRExperimental vLLM support, router hints, cross-node DRAM sharing, local disk caching, resiliency, and NIXL transfersNo native Cartridge manifest, composition contract, or LMCache protocol compatibility.

Measured scope and detailed numbers live on the linked result pages below. “Tested” here describes a component, not an assertion that every row has been integrated with every other row.

LMCache offers the shortest first outlet

The token-dropping SDK already connects a custom KV transformation to actual decoding. That makes it a practical way to test Cartridge-like compacted caches while the native artifact contract matures.

What the API provides

Real cache surgery

The callback receives cached tensors and token IDs, then returns a new KV tensor and the IDs for its new sequence. The SDK checks that the token count matches the KV length, retains the unaligned tail, and stores the edited prefix for inference.

Review the LMCache token-dropping examples →

What remains outside the API

Artifact identity and native loading

A trained Cartridge has no required source-token sequence. It needs an opaque identity, virtual positions, model and layout compatibility, separate K/V dtypes, ordered composition, and a route that can load it before any document prefill.

Inspect the SnapKV implementation →

SnapKV demonstrates the identity problem clearly. The retained positions depend on the query tensor. Reusing that edited cache for another query under a prefix-only key can serve the wrong selection. Either include the query and transformation manifest in the identity, or keep the full shared prefix immutable and perform query-aware selection at request time.

LMCache and KVCR are data-plane choices, not the router

Both can sit below request-to-Cartridge selection. They overlap in placement and movement, but expose different APIs and do not become interoperable merely because both can use NIXL.

LMCACHE

Already supplies KV storage, tiering, engine connectors, a Cartridge storage-plugin path, and the transformation SDK. A native Cartridge API would add opaque artifact lookup and virtual-prefix metadata beside ordinary token-keyed caches.

KV CACHE RUNNER (KVCR)

Pairs naturally with a KV-aware router: the router holds global placement, KVCR manages local DRAM and disk, vLLM owns framework memory, and NIXL executes transfers. Its public release is experimental and currently defines blocks, not trained Cartridge objects.

NIXL is transport, not the object model. KVCR's current G3 tier uses bounded files, fixed slots, and NIXL FILE registrations. A conventional-NVMe backend can extend that path, but raw block storage needs an explicit descriptor adapter, bounded device regions, durable recovery metadata, and device-identity checks. NIXL alone does not define Cartridge selection, ordering, or compatibility.

KV Cache Runner source and design

Make the artifact portable before making it clever

A serving request should name an ordered set of immutable artifacts. Locations and residency can change without changing what the request means.

Semantic identity

Stable

Artifact, version, ordered composition, model geometry, and payload digest define what the request means.

Physical placement

Mutable

HBM, host DRAM, local NVMe, and remote replicas can change while the artifact ID remains stable.

Execution policy

Explicit

Fallback, load deadline, quantization support, and incompatibility behavior must be visible rather than silently guessed.

A narrow path to ecosystem support

Each step produces a useful deployable slice and an observable acceptance test.

Ship the prompt-derived LMCache prototype.Use the existing transformation SDK and real-engine replay to validate orchestration, variable lengths, rerotation, and derived-cache identity on one query-independent Cartridge-like transform.
Freeze a native Cartridge manifest.Represent opaque IDs, exact virtual lengths, model geometry, separate K/V dtypes, checksums, and immutable versions independently of LMCache or KVCR.
Carry an artifact hint through the router and vLLM.Resolve one Cartridge ID first. Require correct allocation, position offsets, concurrent isolation, fallback, and warm/cold accounting.
Attach LMCache or KVCR below the same engine contract.Prove that changing the data plane does not change request semantics. Measure cold fetch, warm residency, eviction, and failure recovery.
Promote K16/V8 to a manifest-level representation.Serialize and transfer BF16 keys and FP8 values separately, then validate quality and the fused prefill/decode path across every tier.
Add ordered CAS composition.Route two or more jointly compatible Cartridges, preserve exact boundaries and visibility, and compare one fused attention computation with ordinary document prefill.
Qualify economics under real reuse.Combine measured training, storage, load, residency, and inference costs with observed popularity and reuse distributions instead of assuming every artifact stays hot.

The remaining gaps are contracts and joins

PORTABLE FORMAT

No shared manifest yet binds logical positions, model geometry, K/V dtype, packing, provenance, and checksums across engines and stores.

VIRTUAL PREFIX API

Prompt-token replacement proves injection, but a trained artifact needs first-class positions that are not claimed to be the request's source tokens.

ORDERED COMPOSITION

Dispatching different requests to different Cartridges is tested. Loading an ordered CAS set into one request is a separate operation.

HETEROGENEOUS FUSION

The current fused kernel handles one shared K16/V8 Cartridge. Multiple lengths, representations, and visibility regions need one online-softmax plan.

INTEROPERABLE STORAGE

LMCache and KVCR have useful but different object and control contracts. An adapter needs explicit ownership, compatibility, and failure semantics.

DEPLOYMENT EVIDENCE

Cold starts, broad workload quality, CUDA graphs, more GPU families, multi-tenant isolation, upgrade compatibility, and full CAS economics remain to be measured.