Memory-Traffic Saturation in Autoregressive Transformer Decode

Decode throughput
has a ceiling.

More KV-cache capacity lets a GPU hold more concurrent sequences. How much faster it serves them depends on where it sits on the throughput curve. The Saturation Law describes that diminishing return; quantization changes both the capacity available and, potentially, the cost of processing each token.

The serving study connects that systems question to numerical quality. Keeping keys at 16 bits and values in FP8 rescues the tested fragile-key model, while reducing conventional KV payload by 25%. Throughput gains depend on the model, context, batch, and executing kernel.

The planning question How much of the extra capacity becomes useful throughput within the latency budget?
EMPIRICAL SATURATION MODEL MEASURED K16/V8 SERVING CAPACITY ≠ THROUGHPUT

Explanation and presentation updated September 15, 2026. Earlier measurements retain their stated model, hardware, and implementation scope.

01 · The Saturation Law

More simultaneous work. A smaller marginal gain.

The law fits aggregate decode throughput as active batch grows. At a fixed context length, throughput rises quickly at first and then flattens. A larger cache can admit more sequences, but it cannot by itself make the GPU process those sequences proportionally faster.

S(B) = Smax · Bγ / (B½γ + Bγ)
B
Active sequences in a decode batch. Requested client concurrency may include queued requests and is a different measurement.
S(B)
Total output tokens per second across that batch, rather than the token rate seen by one user.
Smax
The fitted throughput ceiling for the configuration and context being measured.
B½
The batch that reaches half the fitted ceiling. It describes where diminishing returns become important.
γ
The shape exponent. Fit it from measurements; γ = 1 gives the simple saturating curve illustrated below.

Before the plateau

More active sequences can amortize shared work, including weight access, and improve GPU utilization. Relieving a KV-capacity limit here can unlock substantial throughput.

Near the plateau

Each new sequence still adds KV traffic and computation. More sequences share nearly the same total token rate, so additional capacity buys little throughput and can worsen per-user latency.

This is an empirical Hill fit. It describes a response curve; a good fit does not identify a unique bottleneck or establish that architecture is irrelevant. Model, GPU, context, precision, kernel, parallelism, and scheduling all define the measurement. Fit each relevant context separately. If the measured batches never approach the plateau, the extrapolated ceiling is poorly determined.

What does extra KV capacity buy?

Illustration: γ = 1, unchanged throughput curve, KV-limited batch, enough demand. These are calculated values, not benchmark observations.

For γ = 1, 33⅓% more batch at half the ceiling increases throughput by 14.3%. At 90% of the ceiling the gain is 2.6%.
Starting batchExpanded batch

33.3% more batch → 14.3% more throughput

50.0% → 57.1% of the ceiling. Approximate time per output token rises 16.7%.

The latency estimate uses batch / aggregate throughput for one token per active sequence per decode step. It excludes queueing, prefill, and speculative acceptance. Real serving must measure latency directly.

25% fewer KV bytes can mean 33⅓% more KV capacity. For equally sized K and V tensors, K16/V16 uses four bytes per element pair and K16/V8 uses three. The inverse ratio is 4/3. This is a payload calculation: weights, scale metadata, block fragmentation, workspaces, shared prefixes, and offload buffers still consume memory. Compressed-latent and hybrid caches require their own accounting. None of these ratios establishes numerical quality.

Moving along the curve and changing the curve are different effects. More capacity can increase active batch. A faster attention kernel can raise throughput at the same batch. Reconstruction or transfer work can lower it. Measure fixed-batch performance first, then compare the larger admissible batch under the same quality and latency requirements. Compression may still help at saturation if it reduces the traffic that sets the ceiling; the illustration holds that effect fixed.

The marginal return, expressed mathematically

The local percentage response to batch is d ln S / d ln B = γ · (1 − S/Smax). At γ = 1 and 90% of the ceiling, a small 1% batch increase buys about 0.1% more throughput. For a finite capacity multiplier f at γ = 1, the throughput multiplier is f · (B½ + B) / (B½ + fB).

A useful intuition is step time ≈ a + bB: shared cost a plus per-sequence cost bB. Then aggregate throughput B/(a + bB) has the γ = 1 Hill form. This simplified model explains the shape; it is not a proof that real kernels have only these two costs.

The earlier INT4 speedup fit measures a different quantity

The earlier controlled H100 study fit the same functional form to the speedup ratio of fused INT4 Triton attention against FP16 PyTorch SDPA: 767 decode-mode configurations with B ≥ 4 across 14 models. Its reported parameters were Rmax = 3.75×, B½ = 5.1, γ = 1.32, R² = 0.80. Here the vertical axis is a dimensionless ratio, not aggregate tokens per second.

The fitted curve reaches at least 80% of its asymptotic ratio at B = 16 and at least 95% at B = 64. These coefficients belong to that kernel comparison. They do not predict a 3.75× K16/V8 serving gain, and R² = 0.80 describes fit quality in the observed sample rather than universal predictive accuracy.

Historical fitted INT4 attention ratio: ceiling 3.75×, half-saturation batch 5.1, exponent 1.32.

Only the fitted curve is plotted. No inferred points are labeled as measured observations. See the fused-quantization experiments for the controlled kernel work.

The ideas behind the measurements

These explainers expand the attention loop, hardware scaling, and statistical interpretation used in the study.

Serving evidence

Full-Stack vLLM + FlashInfer Validation — Qwen2.5-7B-Instruct on H100

The asymmetric K16/V8 result now runs end-to-end through the modified vLLM + FlashInfer serving stack (vLLM branch 20260702-k16fp8, fork commit 8a1714108; FlashInfer branch 20260702-asym-k16v8-decode-upstream, fork commit 6dfdc833) with explicit FlashInfer backend selection. The asymmetric path exercises prefill-time V quantization on cache write, paged-cache writes, and an asymmetric FlashInfer decode kernel; prefill computes fresh full-precision K/V. The table reports controlled FP16 / FP8-sym / K16/V8 measurements on the same H100 path: K16/V8 matches FP16 perplexity to reported precision at both 2K and 8K context, and the 200-problem GSM8K screen did not resolve a difference between K16/V8 (90.0%) and FP16 (90.5%) — a screen of this size can detect collapse (symmetric FP8: 2.0%) but cannot establish equivalence.

Config PPL@2K PPL@8K GSM8K (n=200) Smoke tok/s
FP16 baseline 6.997 5.243 90.5% 102.6
FP8 symmetric 214.3 1058.0 2.0% 97.9
Asymmetric K16/V8 6.997 5.243 90.0% 105.8

Symmetric FP8 PPL gets 5× worse from 2K to 8K (214→1058) — the K-fragility phenomenon strengthens at longer context. Asymmetric K16/V8 matches FP16 perplexity to the reported precision at both context lengths. K16 denotes native 16-bit keys: the recipe below uses BF16. Historical FP16 labels are retained for their original comparisons; do not assume every table uses an identical baseline. This is the serving-stack validation reported here; the earlier HuggingFace DynamicCache simulation in the precision-asymmetry section is a separate numerical cross-check.

Tensor parallelism was tested at TP = 1, 2, and 4 on a 4×H100 pod. The reported 48-sequence test found bit-identical prefill NLL and 100% autoregressive token agreement for K16/V8 versus the reference at each tested degree. At TP = 4, the four KV heads shard to one head per rank. These are finite-sample correctness observations, not proof of zero quality cost on all workloads. Fresh full-precision prefill K/V also means prefill agreement alone cannot qualify reads from a quantized cache.

Implementation checks on the reported fork builds
Gate Result
FlashInfer decode, BF16-K / FP8-V rel err 0.0254 vs BF16 ref
FlashInfer prefill, BF16-K / FP8-V rel err 0.0255 vs BF16 ref
vLLM cache writer K bit-exact, V within FP8 noise
K/V aliasing check disjoint storage
Required backend FlashInfer via attention_config
Recipe for the measured fork pair
from vllm import LLM, SamplingParams

llm = LLM(
    model="Qwen/Qwen2.5-7B-Instruct",
    dtype="bfloat16",
    kv_cache_dtype=("auto", "fp8_e4m3"),         # K native, V FP8
    attention_config={"backend": "FLASHINFER"},  # required: auto picks FlashAttn
)

The VLLM_ATTENTION_BACKEND environment variable is not honored in this vLLM build — pass attention_config={"backend": "FLASHINFER"} explicitly. Auto-selection picks FlashAttention, which lacks the asymmetric tuple writer.

HEADLINE

K16/V8 — A Qualified Option for Fragile Keys

Storing keys at native 16-bit precision and only quantizing values to FP8 removes FP8 K operand preparation before each tile's QK operation. Account for FP8 V preparation before PV and measure how much the kernel overlaps it. On the standalone CUDA-core FlashInfer decode kernel on the headline Qwen2.5-7B lane, asymmetric K16/V8 is close to the FP16 reference in this measurement and leads the measured symmetric FP8 path; on the SM90 Tensor Core kernel family the sign reverses — symmetric K8/V8 runs 16–19% faster than K16/V8 at bandwidth-bound shapes (see the linked hardware notes). Across the eight-model vLLM batch-32 comparison below, asymmetric is up to 4.2% faster than symmetric FP8 on the models where it leads. K16/V8 provides 1.33× the KV payload capacity of K16/V16, but only two-thirds of K8/V8 capacity. This implementation uses modified vLLM/FlashInfer kernels without changing model weights or retraining. Quality qualification remains necessary.

Conventional equally sized K/V tensors; raw payload ratios only.
LayoutBytes per K/V pairKV capacity vs K16/V16Qualification
K16/V1641.00×Full-precision reference
K16/V831.33×Retains key precision; qualify value compression
K8/V822.00×Qualify both keys and values for the workload
Mechanism — why bytes alone do not predict throughput
Hopper has no native BF16-query by plain-FP8-key QK instruction with a K scale operand. Prepare each active K tile on chip before QK and advance online softmax tile by tile. Compare the unhidden preparation and occupancy cost against the extra K16 HBM byte. Account for V preparation before PV. CDNA 3 provides FP8-by-FP8 MFMA with no scale operands; place the scale in software around MFMA. Use the completed bias-aware KV findings and the hardware notes.
Quality rescue on fragile-key models
NIAH multikey-3 retrieval at 32K tokens on H100 via vLLM + FlashInfer, across three models:
Qwen2.5-7B (fragile)   FP16: 0.84  ·  FP8-sym: 0.00  ·  Asym K16/V8: 0.89
Qwen3-8B (tolerant)   FP16: 0.95  ·  FP8-sym: 0.92  ·  Asym K16/V8: 0.95
Llama-3.1-8B (tolerant)   FP16: 1.00  ·  FP8-sym: 1.00  ·  Asym K16/V8: 1.00
Qwen2.5-7B scores zero under symmetric FP8 in this retrieval test. The tested K16/V8 retrieval score is 0.89 versus 0.84 for the Qwen2.5-7B reference. Without an uncertainty estimate, this difference does not establish superiority. The two other models show no collapse in this test. When one qualified KV-cache scheme must cover these tested robust- and fragile-key models, asymmetric K16/V8 is the preferred option among the evaluated fused FP8 paths: it is the only one here that holds quality on the fragile models while matching the tolerant ones. The later diagnosis helps explain these observations: symmetric FP8-K collapse tracks K-projection bias magnitude (with a distinct partial-RoPE failure mode), so on a known-robust model symmetric FP8-K is quality-admissible and, being denser, can be the better choice. See the FP8 KV-cache failure atlas for the root cause and the per-model fragility map. A fuller comparison of K16/V8 against the broader KV-quantization field — symmetric FP8, pre-bias, sub-8-bit values, and other approaches across quality, capacity, kernel cost, and served goodput — is now available in the KV-cache compression frontier.
Code: github.com/mcgrof/flashinfer · serving branch 20260702-asym-k16v8-decode-upstream (fork commit 6dfdc833; curated decode-only K16/FP8 dtype-split series). vLLM-side plumbing on 20260702-k16fp8 (fork commit 8a1714108). The asym-prefill-refactor-stage dev tip carries a five-commit CUDA template refactor giving independent K and V dtypes in prefill and decode; its mixed-dtype prefill kernels compile and pass fp32-reference audits, but it is not part of the validated serving pair. Integrates with vLLM via LLM(kv_cache_dtype=("auto", "fp8_e4m3"), attention_config={"backend": "FLASHINFER"}).
THROUGHPUT

Per-Model FlashInfer Throughput — Asymmetric K16/V8 vs Symmetric FP8 on H100

vLLM 0.19 + FlashInfer decode throughput measured at batch 32 on H100 SXM5 across eight open-weight models spanning three architectures. The comparison is three-way: FP16 baseline, symmetric FP8, and our asymmetric K16/V8 branch. Asymmetric reaches 1.38× FP16 on Qwen2.5-7B; on that model the FP16 baseline at batch 32 is cache-capacity-limited, and both FP8 configurations clear that limit, so the 1.34–1.38× reflects capacity relief, not kernel speed. The remaining K16/V8 ratios range from 0.871× to 1.029×. In particular, the 0.5B model regresses by about 13%. Weight-dominated traffic can limit the benefit of KV compression, but identifying the cause of a particular regression requires profiling. Symmetric FP8 trails asymmetric on five of the eight models in this fixed-batch comparison and leads on the other three. Measure the cause with K-only and V-only dtype controls, Blackwell transform modes, warm and cold cache states, and profiler traces. Attribute the gap only when the K-preparation signal moves with the latency delta.

Published H100 batch-32 serving throughput. Ratios use the 16-bit baseline in the same row.
Model16-bit KV (tok/s)K8/V8 (tok/s)K16/V8 (tok/s)K16/V8 / baseline
Qwen2.5-0.5B3,312.33,098.92,884.40.871×
Qwen2.5-1.5B2,912.92,741.32,857.40.981×
Qwen2.5-7B2,065.72,763.02,845.01.377×
Mistral-7B2,861.82,915.92,945.41.029×
Llama-3.1-8B2,084.92,085.52,105.21.010×
Qwen3-8B2,467.32,282.82,315.50.938×
Phi-4-14B1,027.61,041.51,018.40.991×
DS-R1-Qwen-7B2,979.52,896.42,866.40.962×
Asymmetric leads symmetric FP8 on five of the eight tested models at fixed batch 32
Where asymmetric leads, the per-model ratio ranges from +1.0% (Llama-3.1-8B) to +4.2% (Qwen2.5-1.5B), with the headline +2.9% on Qwen2.5-7B (1.38× vs 1.34× FP16); it trails symmetric FP8 on Qwen2.5-0.5B (−6.9%), DS-R1-Qwen-7B (−1.0%), and Phi-4-14B (−2.2%). Treat the ordering as the measured result. Identify its cause with the tile-path matrix: K16/V8 removes FP8 K preparation, while K8/V8 saves one HBM byte per K element. Measure overlap, occupancy, cache state, and shape. Note the scope: this is a fixed-batch-32 end-to-end serving-throughput result on the vLLM 0.19 / FlashInfer 0.6.7 stack reported here, whose asymmetric decode routes to the CUDA-core kernel. The Qwen2.5-7B row includes baseline capacity pressure; this is not an isolated kernel-speed comparison. A separate equal-memory test is needed to measure the benefit of admitting a larger batch. Quality remains a separate, model-specific condition: asymmetric is the configuration here that stays quality-safe on fragile-key models — the NIAH 0.00 → 0.89 retrieval rescue on Qwen2.5-7B — while symmetric FP8-K is fine on robust models (see the FP8 KV-cache failure atlas for which models fail and why).
Measurements at batch 32 through vLLM 0.19 + FlashInfer 0.6.7 with our K16/V8 branch applied (20260702-asym-k16v8-decode-upstream, fork commit 6dfdc833; vLLM fork commit 8a1714108); the table above preserves the published values; the reproduction profile selects the harness. Asymmetric K16/V8 uses the vLLM/FlashInfer dtype-split extension described here. Symmetric FP8 maps onto existing FlashInfer paged-cache infrastructure without modification.
R&D

Earlier R&D — Fusion Principle Proof (Custom INT4 Triton Kernel)

The asymmetric FlashInfer result above rests on a structural claim: kernel fusion — dequantizing KV inside the attention loop instead of materializing an intermediate FP16 buffer — is what turns KV compression into real decode throughput. That claim was first established by a custom fused INT4 Triton kernel before this study integrated the asymmetric FlashInfer serving path. It remains on the repository as a controlled experiment and is now demoted to R&D/appendix status; the paper's main result is the FlashInfer path (the headline asymmetric section and the per-model throughput above).

Non-fused INT4 Dequant to an intermediate FP16 buffer, then call standard attention. The temp-buffer write negates the bandwidth savings from reading smaller values. 0.5× (slowdown vs FP16)
Fused INT4 Triton Dequantize K/V inside the attention tile loop — unpacked values live only in registers and never touch global memory. 2.7–4.8× (H100) · 1.6–7.2× (W7900)
Why this stays on the record
The tested unfused path regresses: materializing a temporary full-precision cache adds traffic that consumes the compression benefit. This result motivates fusion; it does not rule out capacity or reuse benefits in other systems that reconstruct data outside the critical decode loop. The fused INT4 Triton kernel proves the mechanism in isolation from FlashInfer's implementation details. Full write-up and kernel source: knlp/docs/fused-quant; appendix of the paper.
Traffic regime

KV capacity and KV bandwidth are different limits.

A cache can fill GPU memory while weight reads still dominate the cost of a decode step. At larger batch or longer context, KV traffic can become a larger share of the work. Grouped-query attention reduces the KV head count and therefore the payload. Model parameter count alone does not determine the crossover.

fKV ≈ B · T · bKV / (W + B · T · bKV + O)

Here T is cached context length, bKV is KV bytes per token across all layers, W is weight traffic per step, and O is other traffic. This is a simplified byte model, not profiler attribution. KV compression has a smaller direct bandwidth opportunity when fKV is small, even if it relieves a capacity limit. Larger B and T also change kernels and execution cost, so confirm the bottleneck with measurements.

This is why fitting only B × T can obscure the deployment question: B controls how many sequences share a decode step, while T controls how much history each sequence brings. Hold T fixed when studying batch saturation, then evaluate the relevant context distribution in serving.

KV ASYM

KV Precision Asymmetry — Keys and Values Are Not the Same

Earlier activation round-trip quantization probes reveal a precision asymmetry: value-only INT4 incurs much less perplexity damage than key-side INT4 in the shown probes. These numerical probes do not demonstrate a native INT4 serving cache or universal INT4 quality safety. Keys exhibit model-dependent precision floors. On Qwen2.5-7B, compressing keys to INT4 (with INT8 values) causes catastrophic collapse — 17,681% PPL increase, 35,000:1 sensitivity asymmetry. On Mistral-7B, the same configuration causes only +1.23% PPL with 0.98 token agreement. The earlier Qwen precision screen placed a sharp degradation between INT7 and INT6 under its quantizer and scaling protocol; that threshold is not a general bit-width rule.

SENSITIVE Qwen2.5-7B

K / V formatΔPPLagreement
KFP16/VFP16—1.00
KFP16/VINT4+0.51%0.94
KINT8/VINT4+2.17%0.93
KINT4/VINT8+17681%0.15
KINT4/VINT4+18305%0.14

TOLERANT Mistral-7B

K / V formatΔPPLagreement
KFP16/VFP16—1.00
KFP16/VINT4+0.44%0.99
KINT8/VINT4+0.21%0.99
KINT4/VINT8+1.23%0.98
KINT4/VINT4+0.87%0.99
Why architectural predictors fail
The tested summary features did not significantly predict key quantization sensitivity in this sample: GQA ratio (ρ=−0.18), RoPE θ (ρ=0.01), attention entropy (r=0.29 at 6 models, down from r=0.89 at 3). Fisher information and covariance-based signals show no significant correlation (Spearman ρ < 0.2, p > 0.3). This negative correlation result does not prove architectural independence. The later failure atlas identifies specific key-bias and partial-RoPE mechanisms. See Spearman ρ explainer.
An earlier calibration screen, with limited scope
Run 5 calibration prompts through INT8 and INT6 key configs (with INT4 values). Compute mean logit error ratio INT6/INT8. Threshold τ = 3.0 classified 13/13 evaluable models correctly in the reported screen. Qwen family: ratio 5.07–5.40 (flag as sensitive). Other tested models: ratio < 2.2. This is evidence for a screening heuristic, not a safety guarantee for unseen models, contexts, formats, or workloads.
Scale attenuation. Qwen key sensitivity attenuates with model scale: catastrophic at 7B (>10% PPL), +49% at 32B, +1.55% at 72B — below practical threshold. Some earlier larger-model INT4 results used weight-quantization proxies and do not establish cache-quantization safety. The direct 72B FP8 activation experiment below is a separate test.
Asymmetric FP16-K / FP8-V rescues fragile-key models without calibration
On Qwen2.5-7B the same fragility observed under INT4 keys reappears under symmetric FP8: WikiText-2 perplexity jumps from 5.2 (FP16) to 774–1046 at 2–16K context under uncalibrated symmetric FP8. Keeping K at FP16 and only quantizing V to FP8 recovers ppl to 5.23–6.07 across the same context range, GSM8K accuracy from 0.0% back to 84.5%, and MMLU from 48.9% back to 76.2% — the screen resolves no difference from the FP16 baseline. No calibration or per-model tuning was required in this experiment. The separate model-level qualification policy determines deployment. See the headline asymmetric section for throughput data and the hardware notes for the separate quality, representation, and hardware execution constraints.
The tested static per-tensor FP8 calibration fails on Qwen2.5-7B
A common reviewer concern is that the symmetric-FP8 baseline is a strawman vs the standard production calibration path. We tested directly with llm-compressor (the KV-quantization calibration tool used in this experiment), 512-sample WikiText-2 calibration set at T=2048, per-tensor static FP8 e4m3 scales loaded by vLLM. On Qwen2.5-7B the result is worse than uncalibrated symmetric FP8: WikiText-2 perplexity jumps to 9.6×105 with GSM8K accuracy at 0%, vs uncalibrated FP8's PPL ≈ 157 / GSM8K 0.5%. The mechanism is the dual of dynamic-calibration failure: a per-tensor scale wide enough to absorb Qwen's K outliers leaves typical channel values rounded to a coarser grid than the unconditional default, destroying bulk precision in exchange for retaining a handful of outlier samples. Asymmetric K16/V8 under the identical protocol yields PPL 7.50 matching FP16 to four decimals and GSM8K 84.5% — an 84.5-point absolute GSM8K improvement over static-calibrated FP8 (84.5% vs 0%). This rejects that per-tensor static calibration protocol. It does not rule out different scale granularities or bias-aware representations; the later pre-bias work repairs retrieval but fails its serving-performance gate.
Hybrid linear + full-attention models — Qwen3.6-27B (Qwen3.5 family)
Qwen3.6-27B uses Qwen3_5ForConditionalGeneration — a hybrid architecture where every fourth layer carries a true KV cache (the rest are linear attention). At T=2048: FP16 PPL 7.358, FP8-sym PPL 7.347 (drops 10.5 absolute points on GSM8K, 35.5% → 25.0%, n=200 8-shot), asymmetric K16/V8 PPL 7.358 with GSM8K 35.5% — sample-by-sample identical to FP16. Asymmetric is also fastest at evaluation time (24 s for the WikiText-2 sweep vs 79 s FP16, 42 s FP8-sym) due to smaller V cache enabling more concurrency. The measured symmetric path remains slower across batch and context: at T=16,384, symmetric FP8 reaches 0.78×–0.80× of FP16 throughput across B ∈ [4, 32]. Interpret these results with the completed kernel and serving decision rather than treating reduced bytes as a universal speed prediction. These results cover the tested hybrid model and protocol. Compressing the attention-layer KV payload does not reduce every layer's recurrent state by the same ratio.
Large-model FP8 activation quant at 72B (closes proxy gap)
Earlier 27B–72B numbers used hook-based weight-quantization as a proxy for activation sensitivity. We now have a direct activation-quant measurement: Qwen2.5-72B-Instruct (FP8-dynamic weights, ~72 GB on H200) on WikiText-103 at T=2048 over 262K tokens. FP16-KV PPL 4.337, FP8-sym PPL 4.355 (+0.41%), asymmetric K16/V8 PPL 4.337 — NLL identical to FP16 to six decimals. The reported aggregate NLL matches at this precision in this 72B activation test. This does not establish identical token distributions or long-context quality, and it does not validate every earlier weight-quantization proxy.
Consistent with KIVI (Liu et al. 2024) which observes channel-wise outlier structure differences between keys and values. Our finding extends from distribution asymmetry to minimum viable bit-width asymmetry. KIVI: arXiv:2402.02750
Hardware scope

Carry the method across GPUs. Refit the curve.

The wider study includes W7900, A100, H100, B200, H200, and MI300X experiments, with different subsets used for memory-traffic characterization, long-context capacity, and dtype comparisons. They are not one matched end-to-end serving matrix. The H100 serving tables above and standalone attention microbenchmarks measure different scopes; their token rates should not share an unlabeled leaderboard.

Use sustained bandwidth and measured latency at the target shape. Peak HBM bandwidth alone does not determine decode speed. Kernel dispatch, cache residency, launch overhead, parallelism, and the memory traffic generated by the model all matter. A device's tested context limit is also specific to the model and configuration, rather than a universal GPU limit.

SPECULATION

Speculation × Quantization — Sub- and Super-Multiplicative Composition

Speculative decoding does not relax bandwidth constraints. The verification step remains a full attention pass over the entire KV cache. Composition with KV quantization is: sub-multiplicative (ρ ≈ 0.62) for aggressively grouped models (Qwen, 4 KV heads) where KV is a small fraction of total bandwidth, and super-multiplicative (ρ up to 1.95) at long context for models with more KV heads (Llama, 8 KV heads) where KV becomes the bottleneck.

ρ = Scombined / (Squant × Sspec)    ρ < 1.0 sub-mult · ρ = 1.0 mult · ρ > 1.0 super-mult
The tested n-gram setup loses benefit at higher batch
At B=1, acceptance ≥ 0.81 through 8K context. At B=16, acceptance is already below 0.65 at 2K context for two of three models. These observations apply to this n-gram workload and scheduler. Batch alone does not determine draft acceptance, and other draft methods need their own measurements.
Tree verification paradox
Tree speculation helps at low acceptance (+21% at α=0.3) but hurts at high acceptance (−17% at α=0.9). This comparison is specific to the tested draft and verification setup; it is not a universal batch threshold for choosing tree or linear speculation.
Offloading and disaggregation

Which bytes cross the link, and how often?

An attention kernel rereading resident KV from GPU memory is a different workload from restoring a cached prefix from NVMe. A third workload transfers prefilled KV to a decode worker. Each has its own transfer frequency, deadline, and concurrency. A per-step bandwidth bound cannot be applied to all three as an SSD requirement.

WorkloadWhen KV movesWhat sets the requirement
Dense decode with streaming overflowNonresident KV needed by attention must be supplied repeatedly.Bytes read per step, active batch, and the decode-step latency budget.
Prefix-cache reload or request resumeMissing KV is restored before the dependent computation; resident reuse avoids repeated storage reads.Miss or resume rate, restored bytes, available overlap, and time-to-first-token or resume latency.
Prefill/decode disaggregationA handoff transfers the required KV between workers.Handoff rate, payload, topology, and transfer deadline. Storage is optional.

A bandwidth example with explicit assumptions

Take the Qwen2.5-7B cache geometry: 28 layers, four KV heads, and head dimension 128. K16/V16 payload is 28 × 4 × 128 × 4 = 56 KiB per cached token. For a hypothetical 131,072-token cache and one sequence, that is 7 GiB. This is byte accounting at an assumed context, not a claim of quality or successful serving at that length.

Calculated payload bandwidth, decimal GB/s; full cache transferred, no metadata or protocol overhead.
RepresentationCache payloadFull reread every 10 msOne reload within 1 s
K16/V167 GiB751.6 GB/s7.52 GB/s
K16/V85.25 GiB563.7 GB/s5.64 GB/s
K8/V8, if qualified3.5 GiB375.8 GB/s3.76 GB/s

The requirement is transferred bytes divided by the available time. Concurrent full rereads multiply the demand; transferring only an offloaded fraction reduces it. A reload service must also sustain its ongoing arrival rate. Pipelining can hide some latency, but it cannot eliminate the bytes or exceed the sustained capacity of the shared path.

What the Saturation Law contributes
The law estimates how much useful decode throughput extra residency could unlock. Offloading can also save prefill recomputation and improve reuse even when the decoder is near its plateau. Transfer overhead and GPU contention can change the curve, so measure the decoder again with offloading enabled.

Prefill/decode disaggregation does not force NVMe offload. Workers can exchange KV through a memory transport. A storage-mediated handoff is an additional design. Direct GPU–NVMe I/O changes the data path and potential host overhead; it does not remove transfer deadlines or create an automatic throughput gain. This page's serving results do not benchmark that path.

GPU pages, transfer groups, and NVMe commands

An attention block's token count, a connector's transfer grouping, and an SSD command's byte size are separate quantities. Packing, coalescing, scatter/gather layout, and command splitting connect them. The same logical KV object can produce different I/O-size and outstanding-I/O distributions.

For storage projections, record cache misses and evictions, bytes per event, transfer deadlines, effective GPU and storage dtypes, packing cost, submitted command sizes, completion latency, and queue depth at the layer being measured. The Saturation Law alone cannot select NVMe maximum transfer size (MDTS), queue depth, SSD count, or predict energy savings.

FUTURE

Future Directions — Reduce KV Access, Not Just Compress It

The paper's closing argument: compression and kernel engineering are the current best path, but future architectures should address the KV bandwidth bottleneck more directly — constraining KV access proportional to available memory bandwidth rather than compressing the entire cache. Two concrete external directions that motivate this:

MXFP8 on Blackwell
Block-scaled FP8 needs a new quality and execution comparison
K16/V8 protects key precision and changes operand preparation. New hardware can change that execution tradeoff, but does not automatically remove model-level key sensitivity. Blackwell adds block-scaled MXFP8 with hardware consumption of per-block scales for narrow Q and K. FlashInfer contains an SM100 block-scaled fused-attention implementation. Validate its paged-cache layout, model shapes, numerical policy, and serving dispatch against plain E4M3 transformation and K16/V8.
Production R&D
What Remains to Ship
Upstream integration: Land the asymmetric allocator, cache writer, decode kernel, serialization, and API path in the serving projects that own them.

Capacity validation: Compare the representation-aware byte model with allocator telemetry, fragmentation, prefix caching, sliding windows, offload, and disaggregated serving.

Runtime evidence: Have engines attest the effective key/value dtype and carry it through benchmark provenance. Extend model, hardware, tensor-parallel, cache-lifecycle, and long-context coverage without reopening the closed pre-bias kernel line.
Current model-level policy
Use symmetric K8/V8 only after the model and workload pass quality qualification. Keep K at 16 bits when that key quantization fails and K16/V8 passes; retain the full-precision reference when neither tier is qualified. A family name or attention_bias=false is insufficient evidence. The bias-aware policy records the closed pre-bias kernel decision: retrieval was repaired, but the tested equal-memory serving comparison returned 0.654× request throughput and 2.056× p95 request latency relative to K16/V8. That implementation is not the serving default. Trained cartridges require their own qualification and matched sharing baseline; ordinary-generation scores do not establish cartridge quality.
REP

Reproduce — Companion Repos and Defconfig Workflow

The full-stack vLLM + FlashInfer asymmetric K16/V8 result depends on three modified serving-stack repos, all public on GitHub. The branches below are the dated upstream-integration-targeting line used by this page: the vLLM K16/V8 serving branch and the decode-only asymmetric FlashInfer series curated for the upstream PR ladder. They serve asym end to end because prefill computes fresh full-precision K/V and only the cache is stored K16/V8, so a decode-only asymmetric kernel suffices.

Repo Branch What it contains
mcgrof/vllm 20260702-k16fp8 (fork commit 8a1714108) Dated fork implementation: K16/V8 cache with a decode-only asym read (prefill stays full-precision), tuple K/V, loud-fail on asym misroute
mcgrof/flashinfer 20260702-asym-k16v8-decode-upstream (fork commit 6dfdc833) Dated fork implementation: curated decode-only K16/FP8 dtype-split series, rebased on upstream main, GPU-gated for the upstream PR ladder
mcgrof/LMCache asymmetric-kv-codec K16/V8 codec, split-tier placement, serde, 74 CPU unit tests
Planning and result provenance: aiconfigurator K16/V8 byte accounting v2 prices conventional K16/V8 caches without claiming unmeasured latency and rejects compressed-latent layouts. InferenceX KV dtype provenance v2 records requested key/value dtypes without retaining raw server command lines and requires runtime attestation before calling a dtype effective.

The knlp defconfig system selects and runs the reproduction workflow. Review its configuration before building:

git clone https://github.com/mcgrof/knlp.git && cd knlp
make defconfig-decode-upstream   # Dated fork profile; verify resolved revisions
make

defconfig-decode-upstream selects the dated branches above. Branch names can move; record and verify the resolved commit IDs before treating a run as a reproduction. The reproduction profiles are: defconfig-decode (core full-stack quality battery + standalone FlashInfer gates + LMCache codec checks), defconfig-decode-sat (saturation model + Hill fit), and defconfig-decode-full (everything possible from the paper, with structured skip reports when cross-GPU hardware is missing). These profiles select git refs, which may be branches rather than immutable commits, clone the companion forks into the parent directory, build the modified serving stack, run the configured stages, and write machine-readable artifacts to results/decode/<run_id>/. Local JSONL telemetry is canonical; W&B and trackerio are optional mirrors. See docs/reproduce/paper-memory-decode.md for the workflow and hardware requirements. These are fork-specific research recipes, not instructions for an unmodified upstream installation. Fresh prefill, cached-prefix prefill, and decode require separate correctness coverage; full-precision fresh prefill alone does not validate mixed-dtype prefix reads.