Cartridges at Scale

An independent reproduction on a frozen Qwen3-8B over the LongHealth medical question-answering benchmark. With Qwen3's shipped sampler restored, the faithful isolated recipe averages 0.580 across five H100-trained cartridges at the paper's 2,048-token completion cap. This page reports the isolated-cartridge results and controls.

Cartridges at Scale (arXiv:2606.04557) · the base Cartridges technique · CAS extensions and R&D · asymmetric cartridge quantization · deployment ecosystem roadmap · harness

0.39
No context
(paper 0.375)
0.855
Full document
(paper 0.874)
0.86
Cartridge-path
ceiling
0.580
Five cartridges
paper cap
0.683
Patient 02
paper cap
0.617
Five cartridges
8,192-token diagnostic
1. What a cartridge is, and what CAS adds 2. What is validated 3. The faithful isolated recipe 4. Five-patient result 5. Completion-cap diagnostic 6. Training-run variation 7. The cartridge path is lossless 8. Training stability 9. Code and reproduction 10. Serving

1. What a cartridge is, and what CAS adds

A cartridge is a trained KV cache. Train it once per document offline, inject it into the model's KV cache at serving time, and skip prefill. The base technique trains that prefix by self-study context distillation: a teacher answers synthetic questions about the document, and the cartridge is trained to match the teacher's next-token distribution. One document, one cartridge, one prefix at serving time.

Cartridges at Scale keeps that per-document cartridge and adds four things:

2. What is validated

The validated result is the isolated-cartridge reproduction. Accuracy uses the CAS Table-15 LongHealth protocol: option-text prompting, thinking enabled, temperature 0.6, a 2048-token generation budget, fuzzy option matching, all 20 questions for each patient, and three generation runs per trained cartridge.

MeasurementOursPaperStatus
No-context baseline (question only)0.3900.375matched anchor
Full document in context0.8550.874matched anchor
Full-document KV through cartridge execution0.860—path control
Isolated cartridges, five-patient H100 mean0.5800.736corrected sampler
Patient 02, H100-trained cartridge0.683 ± 0.029—corrected sampler, three runs
Five-patient mean with an 8,192-token cap0.617—length diagnostic

The no-context and full-document anchors reproduce the paper within about two percentage points. The exact full-document KV control proves that the alternate cache path itself does not explain the learned-cartridge gap. The five-patient number is a mean over five separately trained H100 cartridges. Each patient score averages three sampled generations of the same trained cartridge; those three generations do not measure training-to-training variation.

This reproduction measures isolated-cartridge training only. It does not include cartridge co-loading or mixed-visibility joint-training results.

3. The faithful isolated recipe

Holding the patient and data fixed, replacing the small-batch optimizer regime with the paper's global-batch-128 schedule raised patient 02 from 0.55 to 0.65 ± 0.05. Batch size, learning rate, and schedule changed together, so this comparison establishes the value of the complete regime rather than attributing the gain to one component.

QuantityFaithful setting
ModelQwen3-8B, frozen BF16 weights
Training dataabout 4,400 self-study conversations per patient; teacher top-20 token distributions
Cartridge lengthceil(document tokens ÷ 20); 632 positions for patient 02
Patient 02 layoutone frozen sink + 631 trainable KV positions
InitializationKV of the document's first p tokens under the system template
Optimization80 epochs, global batch 128, Adam, BF16 cartridge
Learning ratelinear warmup from 0.002 to 0.1 over 200 steps, then linear decay
Schedule horizon5,000 steps; epochs stop first at roughly 800–1,130 steps by patient
EvaluationTable-15 generation, 20 questions, three runs, 2,048-token cap

The schedule horizon is not the number of optimizer updates. For patient 02, 80 epochs end at roughly 1,000 updates while the learning rate is still about 0.083. The harness now stops explicitly at a requested step; the upstream trainer's max_optimizer_steps hook saves a checkpoint but does not otherwise stop training.

What the sink is. Position 0 is the frozen KV entry for the first system-template token, Qwen's <|im_start|>. Queries can place otherwise unhelpful attention mass on this stable slot instead of document-bearing slots. The model does attend to it; dropping it produces degenerate results.

4. Five-patient isolated result

Five separate cartridges were trained with the faithful H100 recipe. The 2,048-token column follows the paper's completion limit. The 8,192-token column re-evaluates the same cartridge and questions with enough room for every answer to finish. Both columns use Qwen3's shipped top-k 20 and top-p 0.95 sampler.

Patient2,048-token accuracyCap hits8,192-token accuracyChange
010.4500 / 600.450+0.000
020.6834 / 600.667−0.017
030.60021 / 600.783+0.183
050.6176 / 600.617+0.000
060.5505 / 600.567+0.017
Mean0.58036 / 3000.617+0.037

The five-patient mean remains below the paper's isolated aggregate of 0.736. The comparison also shows that completion length is not one general explanation for the gap: patient 03 improves sharply, while the other four patients move by at most one answer across their 60 sampled generations.

5. Completion-cap diagnostic

A completion counts as capped only when re-tokenizing its stored text reaches the configured token limit. The evaluator's older truncated field instead recorded whether the model closed its thinking block; treating that field as a cap counter incorrectly marked 17 patient-01 answers even though none reached 2,048 tokens.

Observed behaviorMeasured consequence
Patient 03: 21 of 60 answers cappedAccuracy rises from 0.600 to 0.783 at 8,192 tokens
Patients 02, 05, and 06: 4–6 cappedEach score moves by at most 0.017
Patient 01: no answers cappedAccuracy remains 0.450
Five-patient meanAccuracy rises from 0.580 to 0.617

The longer cap recovers eleven answers for patient 03 and one for patient 06, while patient 02 loses one through ordinary sampling variation. It raises the five-patient mean by 0.037, leaving most of the distance to 0.736 unexplained. The paper synthesizes roughly 40,000 conversations per cartridge versus about 4,400 here; these measurements do not isolate that data-scale difference.

6. Training-run variation

Eight new patient-02 trainings isolate reproducibility from generation noise: two fixed-seed controls and six additional seeds. Every cartridge was evaluated in the same session with three 20-question runs and the corrected sampler.

TrainingSeedFinal reported lossAccuracyPer generation run
New run 1420.30280.7830.70 / 0.80 / 0.85
New run 2420.30280.7830.70 / 0.80 / 0.85
New run 3430.31180.7500.70 / 0.75 / 0.80
New run 4440.30380.7670.80 / 0.80 / 0.70
New run 5450.43330.4000.45 / 0.35 / 0.40
New run 6460.29930.7670.70 / 0.80 / 0.80
New run 7470.30440.8000.80 / 0.85 / 0.75
New run 8480.47880.4000.35 / 0.60 / 0.25

The two seed-42 runs produce byte-identical cartridge files and identical generations. Fixed-seed training is deterministic on this stack. Across the seven distinct seeds, however, five cartridges score 0.750–0.800 and two score 0.400, close to the paper's 0.375 no-context result. The mean is 0.667 with a sample standard deviation of about 0.2. The two low-scoring runs also have clearly higher final training loss, identifying a failed training outcome rather than evaluation noise alone.

Common-session checks of existing cartridges

Existing cartridgeTraining hardwareAccuracyPer generation run
ReferenceH1000.7000.75 / 0.70 / 0.65
Half learning rateH1000.7170.75 / 0.60 / 0.80
Prior same-recipe runH1000.7500.65 / 0.80 / 0.80
A100 runA1000.7330.65 / 0.75 / 0.80

All four existing cartridges fall inside the observed seed range. The A100 cartridge exceeds the reference by two correct answers out of 60 in this session, with overlapping generation spread. The earlier 0.833 score therefore does not establish a hardware effect. Likewise, the learning-rate and prior same-recipe differences are smaller than training-seed variation.

7. The cartridge path is lossless

The cartridge evaluator uses a custom FlexAttention Qwen3 forward, so the ordinary full-document anchor alone cannot validate it. The direct control prefills the complete roughly 12,000-token record, captures every exact K and V tensor, writes that full KV into cartridge format unchanged, and evaluates it through the cartridge path. It scores 0.86, matching ordinary full-document inference at 0.855.

The path has no measurable loss. This rules out positional offsets, RoPE handling, serialization, reconstruction, and cache loading as explanations for the learned-cartridge gap. It does not say that a 632-position learned cartridge is lossless; it says the execution path can carry the full document KV without loss. For patient 02, the full document is 12,628 tokens. This control is distinct from the 631 document-derived trainable positions in the learned-cartridge initialization.

8. Training stability

Training seed controls the failure mode

The seven-seed run finds two distinct outcomes. Five seeds finish near loss 0.30 and score 0.75–0.80; two finish above loss 0.43 and score 0.40. Small single-cartridge ablations cannot be interpreted against that spread. Future comparisons need repeated trainings and a training-loss acceptance check; more documents are required to set a general threshold.

The clipping probe produced an unclipped replicate

The 0.1 peak learning rate uses plain Adam without gradient clipping. A 210-step patient-02 probe logged every gradient norm but did not reproduce the loss excursion that motivated a clipping test. The resulting guard threshold never activated, so the completed run is an ordinary same-recipe replicate, not evidence for or against clipping. That replicate's 0.750 score instead exposed the larger training-run variation described above.

9. Code and reproduction

The public research/cartridges_cas harness contains the synthesis, isolated and joint trainers, Table-15 evaluator, cache-path parity control, prompt-grouped validation split, and checkpoint controls.

Three rules are load-bearing:

Compiled FlexAttention on CUDA sm≥80 gives about a 16× training speedup over the raw path. The target flattener also preserves the full available top-k when cumulative probability does not reach its threshold, instead of silently falling back to top-1.

Five-patient H100 reproduction

cas-paper is the profile for the result on this page. It trains patients 01, 02, 03, 05, and 06, evaluates each cartridge at the 2,048- and 8,192-token caps, and runs exact full-document KV through the cartridge path.

make defconfig-cas-paper
CUDA_VISIBLE_DEVICES=0 \
PAPER_DATA_DIR=/path/to/per_patient \
RECORDS_DIR=/path/to/records \
OUT_DIR=/path/to/cas-output \
CART_ROOT=/path/to/cartridges \
PYTHON=/path/to/python \
HF_HUB_OFFLINE=1 make

Eight-training seed spread

cas-training-spread-h100 launches the two seed-42 controls and seeds 43 through 48 across eight H100s, evaluates all eight artifacts in the same software session, checks repeated-seed byte identity, and writes the summary to OUT_DIR/training_spread/results.json.

make defconfig-cas-training-spread-h100
PAPER_DATA_DIR=/path/to/per_patient \
RECORDS_DIR=/path/to/records \
OUT_DIR=/path/to/cas-output \
CART_ROOT=/path/to/cartridges \
PYTHON=/path/to/python \
HF_HUB_OFFLINE=1 make

The spread profile regenerates the eight new training rows above. The reference, half-learning-rate, prior-run, and A100 comparison rows require their separately archived cartridges. Use cas-paper-regime-a100 for a new single-patient A100 control.

10. Serving

The serving side is already built. A cartridge injects as a KVConnectorBase_V1 plugin into vLLM with no core modifications, and on this record the time-to-first-token drops from 757 ms (re-prefilling the ~12K-token document every query) to 81 ms (loading the 611-token KV prefix), a 677 ms median saving per query at equal decode throughput. The connector, the multi-cartridge router, GPU residency and eviction, and an asymmetric-KV storage codec have all been implemented and tested against vLLM and LMCache.

CAS makes that plumbing broadly useful: a document collection can be compiled once, kept resident under a budget, and selected per request without repeatedly prefilling each retrieved document. Quantization of the stored cartridge is a separate serving R&D result and is intentionally documented separately from this training reproduction.