Decode throughput
has a ceiling.
More KV-cache capacity lets a GPU hold more concurrent sequences. How much faster it serves them depends on where it sits on the throughput curve. The Saturation Law describes that diminishing return; quantization changes both the capacity available and, potentially, the cost of processing each token.
The serving study connects that systems question to numerical quality. Keeping keys at 16 bits and values in FP8 rescues the tested fragile-key model, while reducing conventional KV payload by 25%. Throughput gains depend on the model, context, batch, and executing kernel.
Explanation and presentation updated September 15, 2026. Earlier measurements retain their stated model, hardware, and implementation scope.
More simultaneous work. A smaller marginal gain.
The law fits aggregate decode throughput as active batch grows. At a fixed context length, throughput rises quickly at first and then flattens. A larger cache can admit more sequences, but it cannot by itself make the GPU process those sequences proportionally faster.
- B
- Active sequences in a decode batch. Requested client concurrency may include queued requests and is a different measurement.
- S(B)
- Total output tokens per second across that batch, rather than the token rate seen by one user.
- Smax
- The fitted throughput ceiling for the configuration and context being measured.
- B½
- The batch that reaches half the fitted ceiling. It describes where diminishing returns become important.
- γ
- The shape exponent. Fit it from measurements; γ = 1 gives the simple saturating curve illustrated below.
Before the plateau
More active sequences can amortize shared work, including weight access, and improve GPU utilization. Relieving a KV-capacity limit here can unlock substantial throughput.
Near the plateau
Each new sequence still adds KV traffic and computation. More sequences share nearly the same total token rate, so additional capacity buys little throughput and can worsen per-user latency.
This is an empirical Hill fit. It describes a response curve; a good fit does not identify a unique bottleneck or establish that architecture is irrelevant. Model, GPU, context, precision, kernel, parallelism, and scheduling all define the measurement. Fit each relevant context separately. If the measured batches never approach the plateau, the extrapolated ceiling is poorly determined.
What does extra KV capacity buy?
Illustration: γ = 1, unchanged throughput curve, KV-limited batch, enough demand. These are calculated values, not benchmark observations.
33.3% more batch → 14.3% more throughput
50.0% → 57.1% of the ceiling. Approximate time per output token rises 16.7%.
The latency estimate uses batch / aggregate throughput for one token per active sequence per decode step. It excludes queueing, prefill, and speculative acceptance. Real serving must measure latency directly.
25% fewer KV bytes can mean 33⅓% more KV capacity. For equally sized K and V tensors, K16/V16 uses four bytes per element pair and K16/V8 uses three. The inverse ratio is 4/3. This is a payload calculation: weights, scale metadata, block fragmentation, workspaces, shared prefixes, and offload buffers still consume memory. Compressed-latent and hybrid caches require their own accounting. None of these ratios establishes numerical quality.
Moving along the curve and changing the curve are different effects. More capacity can increase active batch. A faster attention kernel can raise throughput at the same batch. Reconstruction or transfer work can lower it. Measure fixed-batch performance first, then compare the larger admissible batch under the same quality and latency requirements. Compression may still help at saturation if it reduces the traffic that sets the ceiling; the illustration holds that effect fixed.
The marginal return, expressed mathematically
The local percentage response to batch is d ln S / d ln B = γ · (1 − S/Smax). At γ = 1 and 90% of the ceiling, a small 1% batch increase buys about 0.1% more throughput. For a finite capacity multiplier f at γ = 1, the throughput multiplier is f · (B½ + B) / (B½ + fB).
A useful intuition is step time ≈ a + bB: shared cost a plus per-sequence cost bB. Then aggregate throughput B/(a + bB) has the γ = 1 Hill form. This simplified model explains the shape; it is not a proof that real kernels have only these two costs.
The earlier INT4 speedup fit measures a different quantity
The earlier controlled H100 study fit the same functional form to the speedup ratio of fused INT4 Triton attention against FP16 PyTorch SDPA: 767 decode-mode configurations with B ≥ 4 across 14 models. Its reported parameters were Rmax = 3.75×, B½ = 5.1, γ = 1.32, R² = 0.80. Here the vertical axis is a dimensionless ratio, not aggregate tokens per second.
The fitted curve reaches at least 80% of its asymptotic ratio at B = 16 and at least 95% at B = 64. These coefficients belong to that kernel comparison. They do not predict a 3.75× K16/V8 serving gain, and R² = 0.80 describes fit quality in the observed sample rather than universal predictive accuracy.
Only the fitted curve is plotted. No inferred points are labeled as measured observations. See the fused-quantization experiments for the controlled kernel work.
The ideas behind the measurements
These explainers expand the attention loop, hardware scaling, and statistical interpretation used in the study.
Full-Stack vLLM + FlashInfer Validation — Qwen2.5-7B-Instruct on H100
The asymmetric K16/V8 result now runs end-to-end through the modified vLLM + FlashInfer serving stack
(vLLM branch 20260702-k16fp8, fork commit 8a1714108; FlashInfer branch
20260702-asym-k16v8-decode-upstream, fork commit 6dfdc833) with explicit FlashInfer backend
selection. The asymmetric path exercises prefill-time V quantization on cache write, paged-cache writes,
and an asymmetric FlashInfer decode kernel; prefill computes fresh full-precision K/V. The table reports
controlled FP16 / FP8-sym / K16/V8 measurements on the same H100 path: K16/V8 matches FP16 perplexity to
reported precision at both 2K and 8K context, and the 200-problem GSM8K screen did not resolve a difference
between K16/V8 (90.0%) and FP16 (90.5%) — a screen of this size can detect collapse (symmetric FP8: 2.0%)
but cannot establish equivalence.
| Config | PPL@2K | PPL@8K | GSM8K (n=200) | Smoke tok/s |
|---|---|---|---|---|
| FP16 baseline | 6.997 | 5.243 | 90.5% | 102.6 |
| FP8 symmetric | 214.3 | 1058.0 | 2.0% | 97.9 |
| Asymmetric K16/V8 | 6.997 | 5.243 | 90.0% | 105.8 |
Symmetric FP8 PPL gets 5× worse from 2K to 8K (214→1058) — the K-fragility phenomenon strengthens
at longer context. Asymmetric K16/V8 matches FP16 perplexity to the reported precision at both context lengths.
K16 denotes native 16-bit keys: the recipe below uses BF16. Historical FP16 labels are retained for their original comparisons; do not assume every table uses an identical baseline.
This is the serving-stack validation reported here; the earlier HuggingFace DynamicCache simulation in
the precision-asymmetry section is a separate numerical cross-check.
Tensor parallelism was tested at TP = 1, 2, and 4 on a 4×H100 pod. The reported 48-sequence test found bit-identical prefill NLL and 100% autoregressive token agreement for K16/V8 versus the reference at each tested degree. At TP = 4, the four KV heads shard to one head per rank. These are finite-sample correctness observations, not proof of zero quality cost on all workloads. Fresh full-precision prefill K/V also means prefill agreement alone cannot qualify reads from a quantized cache.
| Gate | Result |
|---|---|
| FlashInfer decode, BF16-K / FP8-V | rel err 0.0254 vs BF16 ref |
| FlashInfer prefill, BF16-K / FP8-V | rel err 0.0255 vs BF16 ref |
| vLLM cache writer | K bit-exact, V within FP8 noise |
| K/V aliasing check | disjoint storage |
| Required backend | FlashInfer via attention_config |
from vllm import LLM, SamplingParams
llm = LLM(
model="Qwen/Qwen2.5-7B-Instruct",
dtype="bfloat16",
kv_cache_dtype=("auto", "fp8_e4m3"), # K native, V FP8
attention_config={"backend": "FLASHINFER"}, # required: auto picks FlashAttn
)
The VLLM_ATTENTION_BACKEND environment variable is not honored in this vLLM build —
pass attention_config={"backend": "FLASHINFER"} explicitly. Auto-selection picks FlashAttention,
which lacks the asymmetric tuple writer.
K16/V8 — A Qualified Option for Fragile Keys
Storing keys at native 16-bit precision and only quantizing values to FP8 removes FP8 K operand preparation before each tile's QK operation. Account for FP8 V preparation before PV and measure how much the kernel overlaps it. On the standalone CUDA-core FlashInfer decode kernel on the headline Qwen2.5-7B lane, asymmetric K16/V8 is close to the FP16 reference in this measurement and leads the measured symmetric FP8 path; on the SM90 Tensor Core kernel family the sign reverses — symmetric K8/V8 runs 16–19% faster than K16/V8 at bandwidth-bound shapes (see the linked hardware notes). Across the eight-model vLLM batch-32 comparison below, asymmetric is up to 4.2% faster than symmetric FP8 on the models where it leads. K16/V8 provides 1.33× the KV payload capacity of K16/V16, but only two-thirds of K8/V8 capacity. This implementation uses modified vLLM/FlashInfer kernels without changing model weights or retraining. Quality qualification remains necessary.
| Layout | Bytes per K/V pair | KV capacity vs K16/V16 | Qualification |
|---|---|---|---|
| K16/V16 | 4 | 1.00× | Full-precision reference |
| K16/V8 | 3 | 1.33× | Retains key precision; qualify value compression |
| K8/V8 | 2 | 2.00× | Qualify both keys and values for the workload |
Qwen3-8B (tolerant) FP16: 0.95 · FP8-sym: 0.92 · Asym K16/V8: 0.95
Llama-3.1-8B (tolerant) FP16: 1.00 · FP8-sym: 1.00 · Asym K16/V8: 1.00
20260702-asym-k16v8-decode-upstream
(fork commit 6dfdc833; curated decode-only K16/FP8 dtype-split series). vLLM-side plumbing on
20260702-k16fp8
(fork commit 8a1714108). The
asym-prefill-refactor-stage
dev tip carries a five-commit CUDA template refactor giving independent K and V dtypes in prefill and
decode; its mixed-dtype prefill kernels compile and pass fp32-reference audits, but it is not part of
the validated serving pair.
Integrates with vLLM via LLM(kv_cache_dtype=("auto", "fp8_e4m3"),
attention_config={"backend": "FLASHINFER"}).
Per-Model FlashInfer Throughput — Asymmetric K16/V8 vs Symmetric FP8 on H100
vLLM 0.19 + FlashInfer decode throughput measured at batch 32 on H100 SXM5 across eight open-weight models spanning three architectures. The comparison is three-way: FP16 baseline, symmetric FP8, and our asymmetric K16/V8 branch. Asymmetric reaches 1.38× FP16 on Qwen2.5-7B; on that model the FP16 baseline at batch 32 is cache-capacity-limited, and both FP8 configurations clear that limit, so the 1.34–1.38× reflects capacity relief, not kernel speed. The remaining K16/V8 ratios range from 0.871× to 1.029×. In particular, the 0.5B model regresses by about 13%. Weight-dominated traffic can limit the benefit of KV compression, but identifying the cause of a particular regression requires profiling. Symmetric FP8 trails asymmetric on five of the eight models in this fixed-batch comparison and leads on the other three. Measure the cause with K-only and V-only dtype controls, Blackwell transform modes, warm and cold cache states, and profiler traces. Attribute the gap only when the K-preparation signal moves with the latency delta.
| Model | 16-bit KV (tok/s) | K8/V8 (tok/s) | K16/V8 (tok/s) | K16/V8 / baseline |
|---|---|---|---|---|
| Qwen2.5-0.5B | 3,312.3 | 3,098.9 | 2,884.4 | 0.871× |
| Qwen2.5-1.5B | 2,912.9 | 2,741.3 | 2,857.4 | 0.981× |
| Qwen2.5-7B | 2,065.7 | 2,763.0 | 2,845.0 | 1.377× |
| Mistral-7B | 2,861.8 | 2,915.9 | 2,945.4 | 1.029× |
| Llama-3.1-8B | 2,084.9 | 2,085.5 | 2,105.2 | 1.010× |
| Qwen3-8B | 2,467.3 | 2,282.8 | 2,315.5 | 0.938× |
| Phi-4-14B | 1,027.6 | 1,041.5 | 1,018.4 | 0.991× |
| DS-R1-Qwen-7B | 2,979.5 | 2,896.4 | 2,866.4 | 0.962× |
20260702-asym-k16v8-decode-upstream,
fork commit 6dfdc833; vLLM fork commit 8a1714108); the table above preserves the published values; the reproduction profile selects the harness.
Asymmetric K16/V8 uses the vLLM/FlashInfer dtype-split extension described here. Symmetric FP8 maps onto existing FlashInfer paged-cache infrastructure
without modification.
Earlier R&D — Fusion Principle Proof (Custom INT4 Triton Kernel)
The asymmetric FlashInfer result above rests on a structural claim: kernel fusion — dequantizing KV inside the attention loop instead of materializing an intermediate FP16 buffer — is what turns KV compression into real decode throughput. That claim was first established by a custom fused INT4 Triton kernel before this study integrated the asymmetric FlashInfer serving path. It remains on the repository as a controlled experiment and is now demoted to R&D/appendix status; the paper's main result is the FlashInfer path (the headline asymmetric section and the per-model throughput above).
KV capacity and KV bandwidth are different limits.
A cache can fill GPU memory while weight reads still dominate the cost of a decode step. At larger batch or longer context, KV traffic can become a larger share of the work. Grouped-query attention reduces the KV head count and therefore the payload. Model parameter count alone does not determine the crossover.
Here T is cached context length, bKV is KV bytes per token across all layers, W is weight traffic per step, and O is other traffic. This is a simplified byte model, not profiler attribution. KV compression has a smaller direct bandwidth opportunity when fKV is small, even if it relieves a capacity limit. Larger B and T also change kernels and execution cost, so confirm the bottleneck with measurements.
This is why fitting only B × T can obscure the deployment question: B controls how many sequences share a decode step, while T controls how much history each sequence brings. Hold T fixed when studying batch saturation, then evaluate the relevant context distribution in serving.
KV Precision Asymmetry — Keys and Values Are Not the Same
Earlier activation round-trip quantization probes reveal a precision asymmetry: value-only INT4 incurs much less perplexity damage than key-side INT4 in the shown probes. These numerical probes do not demonstrate a native INT4 serving cache or universal INT4 quality safety. Keys exhibit model-dependent precision floors. On Qwen2.5-7B, compressing keys to INT4 (with INT8 values) causes catastrophic collapse — 17,681% PPL increase, 35,000:1 sensitivity asymmetry. On Mistral-7B, the same configuration causes only +1.23% PPL with 0.98 token agreement. The earlier Qwen precision screen placed a sharp degradation between INT7 and INT6 under its quantizer and scaling protocol; that threshold is not a general bit-width rule.
SENSITIVE Qwen2.5-7B
TOLERANT Mistral-7B
llm-compressor (the KV-quantization calibration tool used in this experiment), 512-sample WikiText-2 calibration set at T=2048, per-tensor static FP8 e4m3 scales loaded
by vLLM. On Qwen2.5-7B the result is worse than uncalibrated symmetric FP8: WikiText-2
perplexity jumps to 9.6×105 with GSM8K accuracy at 0%, vs uncalibrated FP8's
PPL ≈ 157 / GSM8K 0.5%. The mechanism is the dual of dynamic-calibration failure: a per-tensor scale wide
enough to absorb Qwen's K outliers leaves typical channel values rounded to a coarser grid than the
unconditional default, destroying bulk precision in exchange for retaining a handful of outlier samples.
Asymmetric K16/V8 under the identical protocol yields PPL 7.50 matching FP16 to four decimals and GSM8K
84.5% — an 84.5-point absolute GSM8K improvement over static-calibrated FP8 (84.5% vs 0%). This rejects that per-tensor static calibration protocol. It does not rule out different scale granularities or bias-aware representations; the later pre-bias work repairs retrieval but fails its serving-performance gate.
Qwen3_5ForConditionalGeneration — a hybrid architecture where every fourth layer
carries a true KV cache (the rest are linear attention). At T=2048: FP16 PPL 7.358, FP8-sym PPL 7.347
(drops 10.5 absolute points on GSM8K, 35.5% → 25.0%, n=200 8-shot), asymmetric K16/V8 PPL 7.358 with
GSM8K 35.5% — sample-by-sample identical to FP16. Asymmetric is also fastest at evaluation time
(24 s for the WikiText-2 sweep vs 79 s FP16, 42 s FP8-sym) due to smaller V cache enabling more concurrency.
The measured symmetric path remains slower across batch and context: at T=16,384, symmetric FP8 reaches
0.78×–0.80× of FP16 throughput across B ∈ [4, 32]. Interpret these results with the completed
kernel and serving decision rather than
treating reduced bytes as a universal speed prediction.
These results cover the tested hybrid model and protocol. Compressing the attention-layer KV payload does not reduce every layer's recurrent state by the same ratio.
Carry the method across GPUs. Refit the curve.
The wider study includes W7900, A100, H100, B200, H200, and MI300X experiments, with different subsets used for memory-traffic characterization, long-context capacity, and dtype comparisons. They are not one matched end-to-end serving matrix. The H100 serving tables above and standalone attention microbenchmarks measure different scopes; their token rates should not share an unlabeled leaderboard.
Use sustained bandwidth and measured latency at the target shape. Peak HBM bandwidth alone does not determine decode speed. Kernel dispatch, cache residency, launch overhead, parallelism, and the memory traffic generated by the model all matter. A device's tested context limit is also specific to the model and configuration, rather than a universal GPU limit.
Speculation × Quantization — Sub- and Super-Multiplicative Composition
Speculative decoding does not relax bandwidth constraints. The verification step remains a full attention pass over the entire KV cache. Composition with KV quantization is: sub-multiplicative (ρ ≈ 0.62) for aggressively grouped models (Qwen, 4 KV heads) where KV is a small fraction of total bandwidth, and super-multiplicative (ρ up to 1.95) at long context for models with more KV heads (Llama, 8 KV heads) where KV becomes the bottleneck.
Which bytes cross the link, and how often?
An attention kernel rereading resident KV from GPU memory is a different workload from restoring a cached prefix from NVMe. A third workload transfers prefilled KV to a decode worker. Each has its own transfer frequency, deadline, and concurrency. A per-step bandwidth bound cannot be applied to all three as an SSD requirement.
| Workload | When KV moves | What sets the requirement |
|---|---|---|
| Dense decode with streaming overflow | Nonresident KV needed by attention must be supplied repeatedly. | Bytes read per step, active batch, and the decode-step latency budget. |
| Prefix-cache reload or request resume | Missing KV is restored before the dependent computation; resident reuse avoids repeated storage reads. | Miss or resume rate, restored bytes, available overlap, and time-to-first-token or resume latency. |
| Prefill/decode disaggregation | A handoff transfers the required KV between workers. | Handoff rate, payload, topology, and transfer deadline. Storage is optional. |
A bandwidth example with explicit assumptions
Take the Qwen2.5-7B cache geometry: 28 layers, four KV heads, and head dimension 128. K16/V16 payload is 28 × 4 × 128 × 4 = 56 KiB per cached token. For a hypothetical 131,072-token cache and one sequence, that is 7 GiB. This is byte accounting at an assumed context, not a claim of quality or successful serving at that length.
| Representation | Cache payload | Full reread every 10 ms | One reload within 1 s |
|---|---|---|---|
| K16/V16 | 7 GiB | 751.6 GB/s | 7.52 GB/s |
| K16/V8 | 5.25 GiB | 563.7 GB/s | 5.64 GB/s |
| K8/V8, if qualified | 3.5 GiB | 375.8 GB/s | 3.76 GB/s |
The requirement is transferred bytes divided by the available time. Concurrent full rereads multiply the demand; transferring only an offloaded fraction reduces it. A reload service must also sustain its ongoing arrival rate. Pipelining can hide some latency, but it cannot eliminate the bytes or exceed the sustained capacity of the shared path.
Prefill/decode disaggregation does not force NVMe offload. Workers can exchange KV through a memory transport. A storage-mediated handoff is an additional design. Direct GPU–NVMe I/O changes the data path and potential host overhead; it does not remove transfer deadlines or create an automatic throughput gain. This page's serving results do not benchmark that path.
GPU pages, transfer groups, and NVMe commands
An attention block's token count, a connector's transfer grouping, and an SSD command's byte size are separate quantities. Packing, coalescing, scatter/gather layout, and command splitting connect them. The same logical KV object can produce different I/O-size and outstanding-I/O distributions.
For storage projections, record cache misses and evictions, bytes per event, transfer deadlines, effective GPU and storage dtypes, packing cost, submitted command sizes, completion latency, and queue depth at the layer being measured. The Saturation Law alone cannot select NVMe maximum transfer size (MDTS), queue depth, SSD count, or predict energy savings.
Future Directions — Reduce KV Access, Not Just Compress It
The paper's closing argument: compression and kernel engineering are the current best path, but future architectures should address the KV bandwidth bottleneck more directly — constraining KV access proportional to available memory bandwidth rather than compressing the entire cache. Two concrete external directions that motivate this:
Capacity validation: Compare the representation-aware byte model with allocator telemetry, fragmentation, prefix caching, sliding windows, offload, and disaggregated serving.
Runtime evidence: Have engines attest the effective key/value dtype and carry it through benchmark provenance. Extend model, hardware, tensor-parallel, cache-lifecycle, and long-context coverage without reopening the closed pre-bias kernel line.
attention_bias=false is insufficient evidence.
The bias-aware policy records the closed pre-bias kernel decision:
retrieval was repaired, but the tested equal-memory serving comparison returned 0.654× request throughput
and 2.056× p95 request latency relative to K16/V8. That implementation is not the serving default.
Trained cartridges require their own qualification
and matched sharing baseline; ordinary-generation scores do not establish cartridge quality.Reproduce — Companion Repos and Defconfig Workflow
The full-stack vLLM + FlashInfer asymmetric K16/V8 result depends on three modified serving-stack repos, all public on GitHub. The branches below are the dated upstream-integration-targeting line used by this page: the vLLM K16/V8 serving branch and the decode-only asymmetric FlashInfer series curated for the upstream PR ladder. They serve asym end to end because prefill computes fresh full-precision K/V and only the cache is stored K16/V8, so a decode-only asymmetric kernel suffices.
| Repo | Branch | What it contains |
|---|---|---|
| mcgrof/vllm | 20260702-k16fp8 (fork commit 8a1714108) |
Dated fork implementation: K16/V8 cache with a decode-only asym read (prefill stays full-precision), tuple K/V, loud-fail on asym misroute |
| mcgrof/flashinfer | 20260702-asym-k16v8-decode-upstream (fork commit 6dfdc833) |
Dated fork implementation: curated decode-only K16/FP8 dtype-split series, rebased on upstream main, GPU-gated for the upstream PR ladder |
| mcgrof/LMCache | asymmetric-kv-codec |
K16/V8 codec, split-tier placement, serde, 74 CPU unit tests |
The knlp defconfig system selects and runs the reproduction workflow. Review its configuration before building:
git clone https://github.com/mcgrof/knlp.git && cd knlp
make defconfig-decode-upstream # Dated fork profile; verify resolved revisions
make
defconfig-decode-upstream selects the dated branches above. Branch names can move; record and verify the resolved commit IDs before treating a run as a reproduction. The reproduction profiles are: defconfig-decode (core full-stack
quality battery + standalone FlashInfer gates + LMCache codec checks),
defconfig-decode-sat (saturation model + Hill fit), and
defconfig-decode-full (everything possible from the paper, with structured skip reports
when cross-GPU hardware is missing). These profiles select git refs, which may be branches rather than immutable commits, clone the companion forks into
the parent directory, build the modified serving stack, run the configured stages, and write
machine-readable artifacts to results/decode/<run_id>/. Local JSONL telemetry is
canonical; W&B and trackerio are optional mirrors. See
docs/reproduce/paper-memory-decode.md
for the workflow and hardware requirements. These are fork-specific research recipes, not instructions for an unmodified upstream installation. Fresh prefill, cached-prefix prefill, and decode require separate correctness coverage; full-precision fresh prefill alone does not validate mixed-dtype prefix reads.