An independent reproduction on a frozen Qwen3-8B over the LongHealth medical question-answering benchmark. With Qwen3's shipped sampler restored, the faithful isolated recipe averages 0.580 across five H100-trained cartridges at the paper's 2,048-token completion cap. This page reports the isolated-cartridge results and controls.
Cartridges at Scale (arXiv:2606.04557) · the base Cartridges technique · CAS extensions and R&D · asymmetric cartridge quantization · deployment ecosystem roadmap · harness
A cartridge is a trained KV cache. Train it once per document offline, inject it into the model's KV cache at serving time, and skip prefill. The base technique trains that prefix by self-study context distillation: a teacher answers synthetic questions about the document, and the cartridge is trained to match the teacher's next-token distribution. One document, one cartridge, one prefix at serving time.
Cartridges at Scale keeps that per-document cartridge and adds four things:
The validated result is the isolated-cartridge reproduction. Accuracy uses the CAS Table-15 LongHealth protocol: option-text prompting, thinking enabled, temperature 0.6, a 2048-token generation budget, fuzzy option matching, all 20 questions for each patient, and three generation runs per trained cartridge.
| Measurement | Ours | Paper | Status |
|---|---|---|---|
| No-context baseline (question only) | 0.390 | 0.375 | matched anchor |
| Full document in context | 0.855 | 0.874 | matched anchor |
| Full-document KV through cartridge execution | 0.860 | — | path control |
| Isolated cartridges, five-patient H100 mean | 0.580 | 0.736 | corrected sampler |
| Patient 02, H100-trained cartridge | 0.683 ± 0.029 | — | corrected sampler, three runs |
| Five-patient mean with an 8,192-token cap | 0.617 | — | length diagnostic |
The no-context and full-document anchors reproduce the paper within about two percentage points. The exact full-document KV control proves that the alternate cache path itself does not explain the learned-cartridge gap. The five-patient number is a mean over five separately trained H100 cartridges. Each patient score averages three sampled generations of the same trained cartridge; those three generations do not measure training-to-training variation.
This reproduction measures isolated-cartridge training only. It does not include cartridge co-loading or mixed-visibility joint-training results.
Holding the patient and data fixed, replacing the small-batch optimizer regime with the paper's global-batch-128 schedule raised patient 02 from 0.55 to 0.65 ± 0.05. Batch size, learning rate, and schedule changed together, so this comparison establishes the value of the complete regime rather than attributing the gain to one component.
| Quantity | Faithful setting |
|---|---|
| Model | Qwen3-8B, frozen BF16 weights |
| Training data | about 4,400 self-study conversations per patient; teacher top-20 token distributions |
| Cartridge length | ceil(document tokens ÷ 20); 632 positions for patient 02 |
| Patient 02 layout | one frozen sink + 631 trainable KV positions |
| Initialization | KV of the document's first p tokens under the system template |
| Optimization | 80 epochs, global batch 128, Adam, BF16 cartridge |
| Learning rate | linear warmup from 0.002 to 0.1 over 200 steps, then linear decay |
| Schedule horizon | 5,000 steps; epochs stop first at roughly 800–1,130 steps by patient |
| Evaluation | Table-15 generation, 20 questions, three runs, 2,048-token cap |
The schedule horizon is not the number of optimizer updates. For patient 02,
80 epochs end at roughly 1,000 updates while the learning rate is still about
0.083. The harness now stops explicitly at a requested step; the upstream
trainer's max_optimizer_steps hook saves a checkpoint but does not
otherwise stop training.
<|im_start|>. Queries can place otherwise unhelpful attention mass
on this stable slot instead of document-bearing slots. The model does attend to
it; dropping it produces degenerate results.Five separate cartridges were trained with the faithful H100 recipe. The 2,048-token column follows the paper's completion limit. The 8,192-token column re-evaluates the same cartridge and questions with enough room for every answer to finish. Both columns use Qwen3's shipped top-k 20 and top-p 0.95 sampler.
| Patient | 2,048-token accuracy | Cap hits | 8,192-token accuracy | Change |
|---|---|---|---|---|
| 01 | 0.450 | 0 / 60 | 0.450 | +0.000 |
| 02 | 0.683 | 4 / 60 | 0.667 | −0.017 |
| 03 | 0.600 | 21 / 60 | 0.783 | +0.183 |
| 05 | 0.617 | 6 / 60 | 0.617 | +0.000 |
| 06 | 0.550 | 5 / 60 | 0.567 | +0.017 |
| Mean | 0.580 | 36 / 300 | 0.617 | +0.037 |
The five-patient mean remains below the paper's isolated aggregate of 0.736. The comparison also shows that completion length is not one general explanation for the gap: patient 03 improves sharply, while the other four patients move by at most one answer across their 60 sampled generations.
A completion counts as capped only when re-tokenizing its stored text reaches
the configured token limit. The evaluator's older truncated field
instead recorded whether the model closed its thinking block; treating that
field as a cap counter incorrectly marked 17 patient-01 answers even though none
reached 2,048 tokens.
| Observed behavior | Measured consequence |
|---|---|
| Patient 03: 21 of 60 answers capped | Accuracy rises from 0.600 to 0.783 at 8,192 tokens |
| Patients 02, 05, and 06: 4–6 capped | Each score moves by at most 0.017 |
| Patient 01: no answers capped | Accuracy remains 0.450 |
| Five-patient mean | Accuracy rises from 0.580 to 0.617 |
The longer cap recovers eleven answers for patient 03 and one for patient 06, while patient 02 loses one through ordinary sampling variation. It raises the five-patient mean by 0.037, leaving most of the distance to 0.736 unexplained. The paper synthesizes roughly 40,000 conversations per cartridge versus about 4,400 here; these measurements do not isolate that data-scale difference.
Eight new patient-02 trainings isolate reproducibility from generation noise: two fixed-seed controls and six additional seeds. Every cartridge was evaluated in the same session with three 20-question runs and the corrected sampler.
| Training | Seed | Final reported loss | Accuracy | Per generation run |
|---|---|---|---|---|
| New run 1 | 42 | 0.3028 | 0.783 | 0.70 / 0.80 / 0.85 |
| New run 2 | 42 | 0.3028 | 0.783 | 0.70 / 0.80 / 0.85 |
| New run 3 | 43 | 0.3118 | 0.750 | 0.70 / 0.75 / 0.80 |
| New run 4 | 44 | 0.3038 | 0.767 | 0.80 / 0.80 / 0.70 |
| New run 5 | 45 | 0.4333 | 0.400 | 0.45 / 0.35 / 0.40 |
| New run 6 | 46 | 0.2993 | 0.767 | 0.70 / 0.80 / 0.80 |
| New run 7 | 47 | 0.3044 | 0.800 | 0.80 / 0.85 / 0.75 |
| New run 8 | 48 | 0.4788 | 0.400 | 0.35 / 0.60 / 0.25 |
The two seed-42 runs produce byte-identical cartridge files and identical generations. Fixed-seed training is deterministic on this stack. Across the seven distinct seeds, however, five cartridges score 0.750–0.800 and two score 0.400, close to the paper's 0.375 no-context result. The mean is 0.667 with a sample standard deviation of about 0.2. The two low-scoring runs also have clearly higher final training loss, identifying a failed training outcome rather than evaluation noise alone.
| Existing cartridge | Training hardware | Accuracy | Per generation run |
|---|---|---|---|
| Reference | H100 | 0.700 | 0.75 / 0.70 / 0.65 |
| Half learning rate | H100 | 0.717 | 0.75 / 0.60 / 0.80 |
| Prior same-recipe run | H100 | 0.750 | 0.65 / 0.80 / 0.80 |
| A100 run | A100 | 0.733 | 0.65 / 0.75 / 0.80 |
All four existing cartridges fall inside the observed seed range. The A100 cartridge exceeds the reference by two correct answers out of 60 in this session, with overlapping generation spread. The earlier 0.833 score therefore does not establish a hardware effect. Likewise, the learning-rate and prior same-recipe differences are smaller than training-seed variation.
The cartridge evaluator uses a custom FlexAttention Qwen3 forward, so the ordinary full-document anchor alone cannot validate it. The direct control prefills the complete roughly 12,000-token record, captures every exact K and V tensor, writes that full KV into cartridge format unchanged, and evaluates it through the cartridge path. It scores 0.86, matching ordinary full-document inference at 0.855.
The seven-seed run finds two distinct outcomes. Five seeds finish near loss 0.30 and score 0.75–0.80; two finish above loss 0.43 and score 0.40. Small single-cartridge ablations cannot be interpreted against that spread. Future comparisons need repeated trainings and a training-loss acceptance check; more documents are required to set a general threshold.
The 0.1 peak learning rate uses plain Adam without gradient clipping. A 210-step patient-02 probe logged every gradient norm but did not reproduce the loss excursion that motivated a clipping test. The resulting guard threshold never activated, so the completed run is an ordinary same-recipe replicate, not evidence for or against clipping. That replicate's 0.750 score instead exposed the larger training-run variation described above.
The public research/cartridges_cas harness contains the synthesis, isolated and joint trainers, Table-15 evaluator, cache-path parity control, prompt-grouped validation split, and checkpoint controls.
Three rules are load-bearing:
Compiled FlexAttention on CUDA sm≥80 gives about a 16× training speedup over the raw path. The target flattener also preserves the full available top-k when cumulative probability does not reach its threshold, instead of silently falling back to top-1.
cas-paper is the profile for the result on this page. It trains
patients 01, 02, 03, 05, and 06, evaluates each cartridge at the 2,048- and
8,192-token caps, and runs exact full-document KV through the cartridge path.
make defconfig-cas-paper CUDA_VISIBLE_DEVICES=0 \ PAPER_DATA_DIR=/path/to/per_patient \ RECORDS_DIR=/path/to/records \ OUT_DIR=/path/to/cas-output \ CART_ROOT=/path/to/cartridges \ PYTHON=/path/to/python \ HF_HUB_OFFLINE=1 make
cas-training-spread-h100 launches the two seed-42 controls and
seeds 43 through 48 across eight H100s, evaluates all eight artifacts in the
same software session, checks repeated-seed byte identity, and writes the
summary to OUT_DIR/training_spread/results.json.
make defconfig-cas-training-spread-h100 PAPER_DATA_DIR=/path/to/per_patient \ RECORDS_DIR=/path/to/records \ OUT_DIR=/path/to/cas-output \ CART_ROOT=/path/to/cartridges \ PYTHON=/path/to/python \ HF_HUB_OFFLINE=1 make
The spread profile regenerates the eight new training rows above. The
reference, half-learning-rate, prior-run, and A100 comparison rows require their
separately archived cartridges. Use cas-paper-regime-a100 for a new
single-patient A100 control.
The serving side is already built. A cartridge injects as a
KVConnectorBase_V1 plugin into vLLM with no core modifications, and on
this record the time-to-first-token drops from 757 ms (re-prefilling the ~12K-token
document every query) to 81 ms (loading the 611-token KV prefix), a 677 ms median
saving per query at equal decode throughput. The connector, the multi-cartridge
router, GPU residency and eviction, and an asymmetric-KV storage codec have all
been implemented and tested against vLLM and LMCache.
CAS makes that plumbing broadly useful: a document collection can be compiled once, kept resident under a budget, and selected per request without repeatedly prefilling each retrieved document. Quantization of the stored cartridge is a separate serving R&D result and is intentionally documented separately from this training reproduction.