KV-Lingo · independent implementation · October 2026
We rebuilt the paper's 4B↔8B translator. Our maps did not reproduce its multi-turn result.
The paper reports that Qwen3-4B and Qwen3-8B can alternate over ten CoQA turns with translated KV caches while staying close to full native re-prefill, regardless of which model starts. We built and trained the same published class of translator for the same named model pair.
The authors did not release their translator weights, exact model revisions, training-row manifest, or loader. We therefore made an independent training stream from the same Nemotron dataset, pinned our model revisions, trained one seed, and tested four saved forward/reverse checkpoint pairs. Every pair failed our rule when 4B started. Three pairs passed both controls when 8B started, but that narrower result is unconfirmed.
What this meansThis is a failed independent reproduction under a non-identical protocol. It is not a refutation of the paper, and it is not evidence that KV translation cannot work.
SAME QWEN3-4B/8B PAIRSAME DATASET, DIFFERENT ROW STREAM1 SEED VS PAPER'S 332 VS PAPER'S 256 COQA DIALOGUES
Paper versus KNLP, without hand-waving
We followed the published architecture and main training recipe closely. We did not recover the authors' exact checkpoint. The unavailable artifacts and the evaluation changes are large enough to prevent an exact reproduction claim.
Item
KV-Lingo paper
This KNLP work
Match
Question
Can cheap learned maps replace target-model re-prefill across many model pairs and tasks?
Can an independently trained Qwen3-4B↔8B pair meet a fixed quality bar when either model starts a repeated-switch conversation?
Related, narrower
Base models
Released Qwen/Qwen3-4B and Qwen/Qwen3-8B. Exact revisions are not reported.
The same named models, pinned to revisions 1cfa9a72 and b968826d.
Names match
Translator
Separate dense linear key and value maps for each corresponding target layer. Capture is before key normalization and RoPE.
The same map shape and capture point. The receiver applies its own key normalization and RoPE.
Method match
Stage 1
FP64 moments over 400 conversations, about 2.8 million prefix positions, followed by an eigensolve at relative tolerance 10−8.
The same solve over 400 independently selected conversations and 3,487,460 prefix positions.
Recipe matches; rows differ
Stage 2
Forward-KL self-distillation for 5,000 updates, batch 8, AdamW, 3×10−5, 5% warmup, cosine decay, no weight decay, gradient clip 1.0.
The same objective and settings. We also retained the 1,000-update maps to test all four 1k/5k direction pairs.
Main recipe match
Training runs
Learning rate selected on 64 held-out conversations. Reported Qwen3-4B/8B linear results average seeds 42, 43, and 44.
Used the paper's selected 3×10−5 rate and a disjoint 64-conversation validation set. Trained seed 42 only.
1 seed vs 3
Released artifacts
No author code, exact row manifest, model revision hashes, or trained translator tensors were available for this study.
Independent implementation and independently trained maps. Public code and aggregate results are released; trained map tensors are not.
Not same checkpoint
Multi-turn test
256 CoQA dialogues, ten alternating turns, both starting models, reasoning off, 128-token answer budget, three translator seeds.
32 exposed CoQA dialogues, ten alternating turns, both starts, reasoning off, stop at newline/EOS or 64 tokens, one seed, four checkpoint pairs.
Core protocol matches
Decision
The plotted translated condition stays close to full re-prefill through turn ten for both starts.
No checkpoint pair passed our complete either-model-first rule. All eight 4B-start comparisons failed.
Result not reproduced
The important distinction: we reproduced the published method class and its main 4B/8B protocol. We did not reproduce the authors' exact training stream or translator weights. “Independent reproduction” is accurate. “Exact checkpoint reproduction” is not.
What “different Nemotron construction” means
We used the same named dataset, not a different corpus:nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. We pinned revision 1a9454ed054b8544503ab8d8c0a519d141a44c5b. The difference is how conversations and turns were selected and packed into the exact prefix/continuation rows used for fitting.
Paper construction
Published shape; unpublished rows
The paper uses a 50/50 mix. Natural prefixes span 10 to 10,000 tokens with an approximately uniform length distribution. Packed prefixes span 12,000 to 16,000 tokens and concatenate natural samples around a focal question and answer. Reasoning-on and reasoning-off samples alternate. The paper chooses an assistant turn and places the switch at the start of its reply. It does not publish the row identities, exact sampler, packing code, or frozen token stream.
KNLP construction
Same shape; deterministic independent rows
We froze a seed-42 stream with 50% natural and 50% packed rows and alternating reasoning status. Natural rows use 32 equal-width length strata from 10 to 10,000 tokens. Packed rows concatenate complete, reasoning-stripped historical conversations before one focal prompt and keep its answer as the continuation. Every packed prefix is 12,000 to 16,000 tokens. Training and validation conversation UUIDs are disjoint.
SAME NOMINAL BUDGET
Both recipes use 400 Stage-1 conversations and 5,000 × 8 = 40,000 Stage-2 examples per direction. Our Stage-1 rows contain 3,487,460 prefix positions; the paper reports about 2.8 million. Equal row counts do not make the token weighting equal.
WHY THIS CAN MATTER
The learned maps see different conversations, selected turns, lengths, and packed neighbors. Any of these can change the cache covariance fit and distillation gradients. We cannot claim that this difference caused the failed result because seed count, exact model revisions, decoding, and evaluation size also differ.
The complete grid, in four facts
The qualification rule is conjunctive. A checkpoint pair succeeds only when both starting schedules pass against both native references.
4B starts
8 / 8 fail
Every comparison that begins with Qwen3-4B fails at least one quality gate.
8B starts
3 / 4 pairs
Three checkpoint pairs pass both references when Qwen3-8B owns the first turn.
Late turn and health
16 / 16 pass
Every turns-6-to-10 gate passes. The treatment and controls contain zero unhealthy rows.
Full rule
0 / 4 pairs
No forward/reverse checkpoint pair passes both starts and both references.
What the translator does
A receiver uses translated cache segments instead of recomputing the whole conversation prefix after a model switch. This study tests output quality. It does not establish an end-to-end latency, cost, or memory saving.
1 · CaptureCollect source keys before key normalization and collect projected values at corresponding layers.
2 · InitializeFit separate token-wise linear key and value maps with closed-form FP64 sufficient statistics.
3 · DistillFreeze both language models and train the maps against native receiver output distributions.
4 · InstallApply receiver key normalization and RoPE, then append only spans missing from the receiver's cache line.
Retained ownership protocol
Two model-specific cache lines, not one universal cache
4B line
4B native8B span translated once4B native
8B line
4B span translated once8B native4B span translated once
Each line retains its earlier segments. A model translates only the missing spans when it catches up. The protocol does not repeatedly translate the entire accumulated prefix.
This is the paper's retained-span design. Both the paper and this work keep one cache line per model and translate only the span that the incoming model lacks. No span passes through more than one translator. This needs about two full cache lines for an equal-shape pair; it is not a universal single cache.
What the multi-turn test actually did
CoQA is Conversational Question Answering. Each item has one passage and a linked sequence of questions. Later questions can depend on earlier answers. We used 32 complete dialogues: 6 CNN, 7 Gutenberg, 7 MCTest, 6 RACE, and 6 Wikipedia. We evaluated the first ten linked questions, which gives 320 scored answers per starting schedule and per arm.
The starting model natively reads the passage and question 1, then generates answer 1. No translated cache is used for this answer.
The other model receives a translated cache for the passage, question 1, and answer 1. It processes question 2 natively and generates answer 2. This is the first translated-cache answer.
On each later switch, the incoming model keeps its old cache line and translates only the question-and-answer span written since it last ran.
Token F1 scores each generated answer against the human CoQA references. F1 does not score the cache tensors. A native re-prefill control supplies the comparison.
What “turn 2” means: 4B starts, writes the turn-1 history, and 8B answers question 2 from the translated 4B cache plus its native processing of question 2. The turn-2 F1 is the score of 8B's generated answer 2. At this point both native controls have the same turn-1 text, so the first-handoff deficit is not caused by later transcript drift.
Our controls and checkpoint grid
The final comparison uses one provider-local grid because a provider migration changed deterministic outputs. Historical and final-provider rows are not mixed.
Evaluation cohort
32 dialogues × 10 linked questions
Both starting models, greedy decoding, reasoning off, one training seed, and five source domains. These 32 dialogues were visible during checkpoint comparison. They are development data, not an untouched test set.
Four map pairs
1,000 or 5,000 updates per direction
Every combination of forward 4B→8B and reverse 8B→4B translators is evaluated. Update count changes the translator, not either language model's size.
Reference A · added control
Same-transcript native re-prefill
The receiver computes native KV over the exact text generated by the translated arm. This is the clean comparison for cache representation because the text is identical and only the incoming KV differs.
Reference B · closest to paper
Independent native alternation
The models alternate as in the treatment, but every receiver natively re-prefills its own evolving history. This is a system-level comparison with full re-prefill. Its transcript can diverge after the first differing answer.
≤ 3 F1 pointsEqual-domain mean deficit overall.
≤ 5 F1 pointsDeficit in every individual domain.
≤ 5 F1 pointsPooled deficit on turns 6 through 10.
≤ 2 pointsUnhealthy-output rate excess.
These are KNLP deployment gates, not thresholds from the paper. Positive deficit means translated-cache output scored worse than the reference. Every starting schedule and both references must pass all four gates. Bootstrap intervals are descriptive; point estimates decide the gate.
All 16 comparisons
Each reference panel contains all four forward/reverse training-step pairs and both starting models. Bars share a −2 to 14 F1-point scale. The text labels carry the decision, so pass and fail do not depend on color.
Reference A · same-transcript native re-prefill
Native KV is recomputed on the translated arm's exact generated history.
Equal-domain overall deficit
Label: point estimate [descriptive 95% conversation-bootstrap interval] · full condition decision. O+D means overall and domain gates failed.
F1k/R1k · 4B+3.68 [1.37, 5.85] · FAIL O+D
F1k/R1k · 8B−0.38 [−2.59, 1.82] · PASS
F1k/R5k · 4B+3.64 [1.69, 5.59] · FAIL O+D
F1k/R5k · 8B−0.61 [−2.77, 1.65] · PASS
F5k/R1k · 4B+3.16 [0.91, 5.47] · FAIL O+D
F5k/R1k · 8B−0.92 [−2.69, 0.72] · PASS
F5k/R5k · 4B+1.92 [−0.84, 4.69] · FAIL domain
F5k/R5k · 8B−0.32 [−2.13, 1.43] · PASS
−203 limit14
Worst-domain deficit
The named domain is the maximum deficit. The five-point line is independent of the overall mean.
F1k/R1k · 4B5.72 · race · FAIL
F1k/R1k · 8B3.33 · gutenberg · PASS
F1k/R5k · 4B5.72 · race · FAIL
F1k/R5k · 8B3.28 · cnn · PASS
F5k/R1k · 4B6.23 · wikipedia · FAIL
F5k/R1k · 8B1.47 · gutenberg · PASS
F5k/R5k · 4B6.55 · race · FAIL
F5k/R5k · 8B2.66 · gutenberg · PASS
−205 limit14
Reference B · independent native alternation
Each native receiver re-prefills an independently evolving alternating history.
Equal-domain overall deficit
Label: point estimate [descriptive 95% conversation-bootstrap interval] · full condition decision.
F1k/R1k · 4B+5.60 [2.47, 8.75] · FAIL O+D
F1k/R1k · 8B−0.69 [−4.04, 2.63] · PASS
F1k/R5k · 4B+4.57 [1.65, 7.54] · FAIL O+D
F1k/R5k · 8B+2.22 [−1.62, 6.50] · FAIL domain
F5k/R1k · 4B+4.41 [1.54, 7.44] · FAIL O+D
F5k/R1k · 8B−1.65 [−4.57, 1.26] · PASS
F5k/R5k · 4B+3.27 [−0.60, 7.21] · FAIL O+D
F5k/R5k · 8B+1.70 [−1.68, 5.63] · PASS
−203 limit14
Worst-domain deficit
An overall value below three can still fail this five-point domain gate.
F1k/R1k · 4B7.22 · gutenberg · FAIL
F1k/R1k · 8B2.79 · cnn · PASS
F1k/R5k · 4B7.60 · wikipedia · FAIL
F1k/R5k · 8B7.08 · cnn · FAIL
F5k/R1k · 4B12.36 · wikipedia · FAIL
F5k/R1k · 8B3.30 · race · PASS
F5k/R5k · 4B9.44 · wikipedia · FAIL
F5k/R5k · 8B4.38 · cnn · PASS
−205 limit14
F = 4B→8B translator training updates. R = 8B→4B translator training updates. Values are rounded labels; the linked JSON retains full precision. All late-turn deficits are within five points. All unhealthy-rate excess values are zero.
The start-model asymmetry is real in this grid, not yet explained
FIRST HANDOFF
When 4B starts and 8B answers question 2, the mean deficit is 14.32 F1 points with the 1,000-update forward map. It falls to 8.10–8.15 points with the 5,000-update forward map. Each slice contains 32 questions per starting schedule. This is exploratory and does not establish a cause.
BEST DEVELOPMENT OBSERVATION
With forward 5,000 and reverse 1,000, 8B-start translated conversations score 67.39 F1. Independent native alternation scores 65.74; same-history native re-prefill scores 66.46. This is a development observation, not a demonstrated quality gain.
Confounded schedule. The first model also determines odd/even question assignment, the first answer, translation order, and later generated history. Large-model initial cache ownership has not been isolated as the cause. Both directions still translate after turn one.
Separate one-switch diagnostic: ours learned, but did not reach native parity
The paper's single-switch CoQA table reports the 4B cache translated once into 8B at 0.746 F1, versus 0.712 for native 8B, on 256 samples and averaged over three translator seeds. Our distinct fixed-history diagnostic used 96 questions: turns 1, 5, and 10 from the same 32 development dialogues. With the 1,000-update 4B→8B map, translated 8B rose from the Stage-1 value of 55.00 to 72.04 F1, but native 8B scored 80.90. The remaining deficit was 8.86 points.
This is evidence that our translator learned useful structure but did not match the paper's reported native-level behavior. It is not a numerical head-to-head: the histories, item set, checkpoint, seed count, and decode contract differ. We did not archive a final 5,000-update run of this isolated 96-question diagnostic. The ten-turn grid above supplies the final-map first-handoff evidence.
Was this a fair test?
Yes for our deployment decision. No for an exact claim that the paper is wrong. Those are different standards.
FAIR FOR THE KNLP GO/NO-GO
All four map pairs used the same pinned models, the same 32 dialogue identities, the same scorer, and a complete provider-local replay. We tested both starting models. The same-transcript control holds generated text fixed and changes only native versus translated KV. The gates were frozen before the final mixed-checkpoint grid was interpreted.
NOT AN EXACT PAPER REPLICATION
The author checkpoints and row loader were unavailable. We used one translator seed instead of three, 32 exposed dialogues instead of 256, a 64-token/newline answer stop instead of the paper's 128-token budget, and our own per-domain gates. The paper also does not report exact base-model revision hashes.
Our thresholds did not create the whole discrepancy. The paper did not define our 3-point overall or 5-point domain limits, so those limits control our pass/fail label. But the raw first 4B→8B handoff with a 5,000-update map still lost about 8.1 F1 points against native re-prefill. For the 5k/5k pair, the ten-turn 4B-start result lost 1.92 points overall against the same-transcript control but 6.55 in its worst domain; against independent native alternation it lost 3.27 overall and 9.44 in its worst domain. The negative result is not only a threshold artifact.
Viability and unresolved issues
UNRESTRICTED SHARING
Arbitrary either-model-first sharing is unqualified for Qwen3-4B/8B under this recipe and these gates.
RESTRICTED 8B-FIRST SCHEDULE
The fixed 8B-first schedule is a narrower service hypothesis. It needs a fresh, untouched confirmation and a full-path value measurement.
GENERALIZATION
One model pair, one seed, 32 exposed development conversations, five domains, and ten turns do not establish arbitrary routing, cross-family transfer, or a long-context guarantee.
STORAGE AND VALUE
Two retained cache lines do not prove one-cache storage savings. This quality study does not measure useful latency, cost, or memory reduction.
Provider migration changed deterministic outputs, so the last study replayed a matched grid under one environment. CPU post-processing corrected receipt-state handling after a deadline refusal. Those operational issues do not invalidate the completed provider-local comparison, but they limit claims about hardware causes and portability.
Multiple checkpoint comparisons increase selection risk. No untouched reserve evaluation was run because no pair passed every development gate. No result here covers a different model family, learned-Cartridge reuse, production readiness, or the paper's reported latency gains.
Source, data, and reproduction boundary
The public code is a clean import of the executed snapshot. Provider launch wrappers and private resource records are excluded. A file-hash manifest records the mapping and packaging changes. The aggregate reproduces the 16 multi-turn decisions. Full inference needs the public models, the pinned data construction, and trained translator maps; the map tensors are not published.
Public source integration: 8bb7f320. Aggregate and guide: 5cf299ed. Aggregate SHA-256: 2051e9abdb85ad49c38872d7caf15d8488d4e3a3795115d363e5355d238d44b1. Raw outputs and trained maps are not released.