Each step produces a useful deployable slice and an observable acceptance test.
Ship the prompt-derived LMCache prototype.Use the existing transformation SDK and real-engine replay to validate orchestration, variable lengths, rerotation, and derived-cache identity on one query-independent Cartridge-like transform.
Freeze a native Cartridge manifest.Represent opaque IDs, exact virtual lengths, model geometry, separate K/V dtypes, checksums, and immutable versions independently of LMCache or KVCR.
Carry an artifact hint through the router and vLLM.Resolve one Cartridge ID first. Require correct allocation, position offsets, concurrent isolation, fallback, and warm/cold accounting.
Attach LMCache or KVCR below the same engine contract.Prove that changing the data plane does not change request semantics. Measure cold fetch, warm residency, eviction, and failure recovery.
Promote K16/V8 to a manifest-level representation.Serialize and transfer BF16 keys and FP8 values separately, then validate quality and the fused prefill/decode path across every tier.
Add ordered CAS composition.Route two or more jointly compatible Cartridges, preserve exact boundaries and visibility, and compare one fused attention computation with ordinary document prefill.
Qualify economics under real reuse.Combine measured training, storage, load, residency, and inference costs with observed popularity and reuse distributions instead of assuming every artifact stays hot.