KV-Lingo · independent implementation · October 2026

We rebuilt the paper's 4B↔8B translator.
Our maps did not reproduce its multi-turn result.

The paper reports that Qwen3-4B and Qwen3-8B can alternate over ten CoQA turns with translated KV caches while staying close to full native re-prefill, regardless of which model starts. We built and trained the same published class of translator for the same named model pair.

The authors did not release their translator weights, exact model revisions, training-row manifest, or loader. We therefore made an independent training stream from the same Nemotron dataset, pinned our model revisions, trained one seed, and tested four saved forward/reverse checkpoint pairs. Every pair failed our rule when 4B started. Three pairs passed both controls when 8B started, but that narrower result is unconfirmed.

What this meansThis is a failed independent reproduction under a non-identical protocol. It is not a refutation of the paper, and it is not evidence that KV translation cannot work.
SAME QWEN3-4B/8B PAIRSAME DATASET, DIFFERENT ROW STREAM1 SEED VS PAPER'S 332 VS PAPER'S 256 COQA DIALOGUES

Paper versus KNLP, without hand-waving

We followed the published architecture and main training recipe closely. We did not recover the authors' exact checkpoint. The unavailable artifacts and the evaluation changes are large enough to prevent an exact reproduction claim.

ItemKV-Lingo paperThis KNLP workMatch
QuestionCan cheap learned maps replace target-model re-prefill across many model pairs and tasks?Can an independently trained Qwen3-4B↔8B pair meet a fixed quality bar when either model starts a repeated-switch conversation?Related, narrower
Base modelsReleased Qwen/Qwen3-4B and Qwen/Qwen3-8B. Exact revisions are not reported.The same named models, pinned to revisions 1cfa9a72 and b968826d.Names match
TranslatorSeparate dense linear key and value maps for each corresponding target layer. Capture is before key normalization and RoPE.The same map shape and capture point. The receiver applies its own key normalization and RoPE.Method match
Stage 1FP64 moments over 400 conversations, about 2.8 million prefix positions, followed by an eigensolve at relative tolerance 10−8.The same solve over 400 independently selected conversations and 3,487,460 prefix positions.Recipe matches; rows differ
Stage 2Forward-KL self-distillation for 5,000 updates, batch 8, AdamW, 3×10−5, 5% warmup, cosine decay, no weight decay, gradient clip 1.0.The same objective and settings. We also retained the 1,000-update maps to test all four 1k/5k direction pairs.Main recipe match
Training runsLearning rate selected on 64 held-out conversations. Reported Qwen3-4B/8B linear results average seeds 42, 43, and 44.Used the paper's selected 3×10−5 rate and a disjoint 64-conversation validation set. Trained seed 42 only.1 seed vs 3
Released artifactsNo author code, exact row manifest, model revision hashes, or trained translator tensors were available for this study.Independent implementation and independently trained maps. Public code and aggregate results are released; trained map tensors are not.Not same checkpoint
Multi-turn test256 CoQA dialogues, ten alternating turns, both starting models, reasoning off, 128-token answer budget, three translator seeds.32 exposed CoQA dialogues, ten alternating turns, both starts, reasoning off, stop at newline/EOS or 64 tokens, one seed, four checkpoint pairs.Core protocol matches
DecisionThe plotted translated condition stays close to full re-prefill through turn ten for both starts.No checkpoint pair passed our complete either-model-first rule. All eight 4B-start comparisons failed.Result not reproduced

The important distinction: we reproduced the published method class and its main 4B/8B protocol. We did not reproduce the authors' exact training stream or translator weights. “Independent reproduction” is accurate. “Exact checkpoint reproduction” is not.

What “different Nemotron construction” means

We used the same named dataset, not a different corpus: nvidia/Nemotron-SFT-Instruction-Following-Chat-v2. We pinned revision 1a9454ed054b8544503ab8d8c0a519d141a44c5b. The difference is how conversations and turns were selected and packed into the exact prefix/continuation rows used for fitting.

Paper construction

Published shape; unpublished rows

The paper uses a 50/50 mix. Natural prefixes span 10 to 10,000 tokens with an approximately uniform length distribution. Packed prefixes span 12,000 to 16,000 tokens and concatenate natural samples around a focal question and answer. Reasoning-on and reasoning-off samples alternate. The paper chooses an assistant turn and places the switch at the start of its reply. It does not publish the row identities, exact sampler, packing code, or frozen token stream.

KNLP construction

Same shape; deterministic independent rows

We froze a seed-42 stream with 50% natural and 50% packed rows and alternating reasoning status. Natural rows use 32 equal-width length strata from 10 to 10,000 tokens. Packed rows concatenate complete, reasoning-stripped historical conversations before one focal prompt and keep its answer as the continuation. Every packed prefix is 12,000 to 16,000 tokens. Training and validation conversation UUIDs are disjoint.

SAME NOMINAL BUDGET

Both recipes use 400 Stage-1 conversations and 5,000 × 8 = 40,000 Stage-2 examples per direction. Our Stage-1 rows contain 3,487,460 prefix positions; the paper reports about 2.8 million. Equal row counts do not make the token weighting equal.

WHY THIS CAN MATTER

The learned maps see different conversations, selected turns, lengths, and packed neighbors. Any of these can change the cache covariance fit and distillation gradients. We cannot claim that this difference caused the failed result because seed count, exact model revisions, decoding, and evaluation size also differ.

The complete grid, in four facts

The qualification rule is conjunctive. A checkpoint pair succeeds only when both starting schedules pass against both native references.

4B starts
8 / 8 fail

Every comparison that begins with Qwen3-4B fails at least one quality gate.

8B starts
3 / 4 pairs

Three checkpoint pairs pass both references when Qwen3-8B owns the first turn.

Late turn and health
16 / 16 pass

Every turns-6-to-10 gate passes. The treatment and controls contain zero unhealthy rows.

Full rule
0 / 4 pairs

No forward/reverse checkpoint pair passes both starts and both references.

What the translator does

A receiver uses translated cache segments instead of recomputing the whole conversation prefix after a model switch. This study tests output quality. It does not establish an end-to-end latency, cost, or memory saving.

1 · CaptureCollect source keys before key normalization and collect projected values at corresponding layers.
2 · InitializeFit separate token-wise linear key and value maps with closed-form FP64 sufficient statistics.
3 · DistillFreeze both language models and train the maps against native receiver output distributions.
4 · InstallApply receiver key normalization and RoPE, then append only spans missing from the receiver's cache line.
Retained ownership protocol

Two model-specific cache lines, not one universal cache

4B-native segment8B-native segmenttranslated segment

Each line retains its earlier segments. A model translates only the missing spans when it catches up. The protocol does not repeatedly translate the entire accumulated prefix.

This is the paper's retained-span design. Both the paper and this work keep one cache line per model and translate only the span that the incoming model lacks. No span passes through more than one translator. This needs about two full cache lines for an equal-shape pair; it is not a universal single cache.

What the multi-turn test actually did

CoQA is Conversational Question Answering. Each item has one passage and a linked sequence of questions. Later questions can depend on earlier answers. We used 32 complete dialogues: 6 CNN, 7 Gutenberg, 7 MCTest, 6 RACE, and 6 Wikipedia. We evaluated the first ten linked questions, which gives 320 scored answers per starting schedule and per arm.

The starting model natively reads the passage and question 1, then generates answer 1. No translated cache is used for this answer.
The other model receives a translated cache for the passage, question 1, and answer 1. It processes question 2 natively and generates answer 2. This is the first translated-cache answer.
On each later switch, the incoming model keeps its old cache line and translates only the question-and-answer span written since it last ran.
Token F1 scores each generated answer against the human CoQA references. F1 does not score the cache tensors. A native re-prefill control supplies the comparison.

What “turn 2” means: 4B starts, writes the turn-1 history, and 8B answers question 2 from the translated 4B cache plus its native processing of question 2. The turn-2 F1 is the score of 8B's generated answer 2. At this point both native controls have the same turn-1 text, so the first-handoff deficit is not caused by later transcript drift.

Our controls and checkpoint grid

The final comparison uses one provider-local grid because a provider migration changed deterministic outputs. Historical and final-provider rows are not mixed.

Evaluation cohort

32 dialogues × 10 linked questions

Both starting models, greedy decoding, reasoning off, one training seed, and five source domains. These 32 dialogues were visible during checkpoint comparison. They are development data, not an untouched test set.

Four map pairs

1,000 or 5,000 updates per direction

Every combination of forward 4B→8B and reverse 8B→4B translators is evaluated. Update count changes the translator, not either language model's size.

Reference A · added control

Same-transcript native re-prefill

The receiver computes native KV over the exact text generated by the translated arm. This is the clean comparison for cache representation because the text is identical and only the incoming KV differs.

Reference B · closest to paper

Independent native alternation

The models alternate as in the treatment, but every receiver natively re-prefills its own evolving history. This is a system-level comparison with full re-prefill. Its transcript can diverge after the first differing answer.

≤ 3 F1 pointsEqual-domain mean deficit overall.
≤ 5 F1 pointsDeficit in every individual domain.
≤ 5 F1 pointsPooled deficit on turns 6 through 10.
≤ 2 pointsUnhealthy-output rate excess.

These are KNLP deployment gates, not thresholds from the paper. Positive deficit means translated-cache output scored worse than the reference. Every starting schedule and both references must pass all four gates. Bootstrap intervals are descriptive; point estimates decide the gate.

All 16 comparisons

Each reference panel contains all four forward/reverse training-step pairs and both starting models. Bars share a −2 to 14 F1-point scale. The text labels carry the decision, so pass and fail do not depend on color.

Reference A · same-transcript native re-prefill

Native KV is recomputed on the translated arm's exact generated history.

Reference B · independent native alternation

Each native receiver re-prefills an independently evolving alternating history.

F = 4B→8B translator training updates. R = 8B→4B translator training updates. Values are rounded labels; the linked JSON retains full precision. All late-turn deficits are within five points. All unhealthy-rate excess values are zero.

The start-model asymmetry is real in this grid, not yet explained

FIRST HANDOFF

When 4B starts and 8B answers question 2, the mean deficit is 14.32 F1 points with the 1,000-update forward map. It falls to 8.10–8.15 points with the 5,000-update forward map. Each slice contains 32 questions per starting schedule. This is exploratory and does not establish a cause.

BEST DEVELOPMENT OBSERVATION

With forward 5,000 and reverse 1,000, 8B-start translated conversations score 67.39 F1. Independent native alternation scores 65.74; same-history native re-prefill scores 66.46. This is a development observation, not a demonstrated quality gain.

Confounded schedule. The first model also determines odd/even question assignment, the first answer, translation order, and later generated history. Large-model initial cache ownership has not been isolated as the cause. Both directions still translate after turn one.

Separate one-switch diagnostic: ours learned, but did not reach native parity

The paper's single-switch CoQA table reports the 4B cache translated once into 8B at 0.746 F1, versus 0.712 for native 8B, on 256 samples and averaged over three translator seeds. Our distinct fixed-history diagnostic used 96 questions: turns 1, 5, and 10 from the same 32 development dialogues. With the 1,000-update 4B→8B map, translated 8B rose from the Stage-1 value of 55.00 to 72.04 F1, but native 8B scored 80.90. The remaining deficit was 8.86 points.

This is evidence that our translator learned useful structure but did not match the paper's reported native-level behavior. It is not a numerical head-to-head: the histories, item set, checkpoint, seed count, and decode contract differ. We did not archive a final 5,000-update run of this isolated 96-question diagnostic. The ten-turn grid above supplies the final-map first-handoff evidence.

Was this a fair test?

Yes for our deployment decision. No for an exact claim that the paper is wrong. Those are different standards.

FAIR FOR THE KNLP GO/NO-GO

All four map pairs used the same pinned models, the same 32 dialogue identities, the same scorer, and a complete provider-local replay. We tested both starting models. The same-transcript control holds generated text fixed and changes only native versus translated KV. The gates were frozen before the final mixed-checkpoint grid was interpreted.

NOT AN EXACT PAPER REPLICATION

The author checkpoints and row loader were unavailable. We used one translator seed instead of three, 32 exposed dialogues instead of 256, a 64-token/newline answer stop instead of the paper's 128-token budget, and our own per-domain gates. The paper also does not report exact base-model revision hashes.

Our thresholds did not create the whole discrepancy. The paper did not define our 3-point overall or 5-point domain limits, so those limits control our pass/fail label. But the raw first 4B→8B handoff with a 5,000-update map still lost about 8.1 F1 points against native re-prefill. For the 5k/5k pair, the ten-turn 4B-start result lost 1.92 points overall against the same-transcript control but 6.55 in its worst domain; against independent native alternation it lost 3.27 overall and 9.44 in its worst domain. The negative result is not only a threshold artifact.

Viability and unresolved issues

UNRESTRICTED SHARING

Arbitrary either-model-first sharing is unqualified for Qwen3-4B/8B under this recipe and these gates.

RESTRICTED 8B-FIRST SCHEDULE

The fixed 8B-first schedule is a narrower service hypothesis. It needs a fresh, untouched confirmation and a full-path value measurement.

GENERALIZATION

One model pair, one seed, 32 exposed development conversations, five domains, and ten turns do not establish arbitrary routing, cross-family transfer, or a long-context guarantee.

STORAGE AND VALUE

Two retained cache lines do not prove one-cache storage savings. This quality study does not measure useful latency, cost, or memory reduction.

Provider migration changed deterministic outputs, so the last study replayed a matched grid under one environment. CPU post-processing corrected receipt-state handling after a deadline refusal. Those operational issues do not invalidate the completed provider-local comparison, but they limit claims about hardware causes and portability.

Multiple checkpoint comparisons increase selection risk. No untouched reserve evaluation was run because no pair passed every development gate. No result here covers a different model family, learned-Cartridge reuse, production readiness, or the paper's reported latency gains.

Source, data, and reproduction boundary

The public code is a clean import of the executed snapshot. Provider launch wrappers and private resource records are excluded. A file-hash manifest records the mapping and packaging changes. The aggregate reproduces the 16 multi-turn decisions. Full inference needs the public models, the pinned data construction, and trained translator maps; the map tensors are not published.

Public source integration: 8bb7f320. Aggregate and guide: 5cf299ed. Aggregate SHA-256: 2051e9abdb85ad49c38872d7caf15d8488d4e3a3795115d363e5355d238d44b1. Raw outputs and trained maps are not released.

Practical statusResearch code merged; checkpoint search closed; unrestricted sharing unqualified; restricted 8B-first use unconfirmed.