Home › Articles › KVarN KV Cache: Implementation and Benchmarks
Tests on Qwen 3.6 27B show how an experimental BeeLlama KVarN preview shifts KV-cache quality down a memory tier: q5 quality at q4 size, q6 at q5, and q8 at q6.
Disclaimer: this is still a narrow benchmark. The KVarN path tested here is a preview implementation in BeeLlama v0.3.2, not a mature optimized runtime. The precision numbers are the meaningful part of the article, but may be subject to change with future updates. The tok/s column is prompt-processing behavior on a very raw preview of this specific fork, not a final verdict on KVarN's decode or generation speed.
The original KV-cache benchmark had a pretty clear sub-6-bit story: normal q4/q5 cache quantization was stronger than TurboQuant at similar practical sizes, while TCQ made the 2-3 bit turbo modes less bad. You could save memory, or you could preserve the output distribution, and the tradeoff was mostly where you expected it to be.
KVarN changes that curve. On the main Q5_K_S 64k comparison, kvarn4-kvarn4 uses 27.9% of the bf16 KV-cache footprint and lands around the q5_0/q5_1 quality tier. kvarn4-kvarn3 uses 24.8%, below ordinary q4_0, and still beats q4_0 on mean KLD. kvarn3-kvarn3 sits near the turbo3 memory class and cuts the turbo3_tcq mean KLD by a third.
That does not make KVarN lossless. The bf16 KLD reference is still far away. But for VRAM-constrained long-context setups, "q5-ish quality at q4-ish memory" is the result people are happy to see for low-bit KV-cache methods. Where TurboQuant promised that kind of shift, KVarN is the first thing in these benches that actually looks like it.
The 5/6/8 follow-up at the end widens that result into a cleaner ladder. kvarn5-kvarn5 behaves like a q6-class row at q5-class memory, while kvarn6-kvarn6 reaches the q8-class ceiling around q6 memory. That is the stronger claim: KVarN is not merely good for its bit width, it often matches the next ordinary cache tier while paying for the cheaper one.
KVarN is a KV-cache quantization method from the paper KVarN: Variance-Normalized KV-Cache Quantization Mitigates Error Accumulation in Reasoning Tasks. The paper's central claim is that current KV-cache quantizers make bad token-scale errors during autoregressive decoding, and those errors accumulate over time, while KVarN reduces the outlier errors that matter most.
The mechanism has two parts. First, KVarN applies a Hadamard rotation in the channel dimension, which spreads outlier energy across coordinates before quantization. That part is familiar from QuaRot-style weight work and from the broader low-bit KV-cache literature. Second, it applies dual-axis variance normalization across both dimensions of a KV tile, so both token and channel scale variation are brought under control before round-to-nearest quantization. The paper describes this as a Sinkhorn-inspired variance-normalization step.
In plain terms: ordinary per-axis quantization can preserve one shape of scale variation and still blow up a few token norms. KVarN spends extra metadata on a second scale so the quantizer can keep those token magnitudes closer to where they were. The paper argues that the worst few percent of quantization errors do disproportionate damage, especially in decode-heavy reasoning tasks, so suppressing those outliers matters more than lowering average reconstruction error everywhere.
BeeLlama is my llama.cpp fork for long-context experiments. The earlier benchmark used it for q6_0, TurboQuant and TCQ cache modes. This KVarN run is the same kind of fork-level experiment, but the KVarN path is much younger. The first pass tested experimental cache types named kvarn2, kvarn3 and kvarn4, plus asymmetric K/V pairs; the follow-up extends the same path to kvarn5, kvarn6 and kvarn8.
The BeeLlama implementation tested here follows the same broad KVarN structure: 128-token groups, 2/3/4-bit payloads, Hadamard rotation, iterative variance normalization, and stored scale/zero-point metadata. The measured sizes in this article are from the actual BeeLlama run logs, which matters because this preview stores auxiliary values as fp16 records, while the paper reports a more optimized accounting with lower auxiliary precision.
KVarN is also not just one more GGUF-style cache quant like q8_0, q5_0, q4_0, or the turbo types. In BeeLlama, kvarn2, kvarn3, and kvarn4 are CLI pseudo-types that select a separate structured KVarN cache backend. The underlying cache tensors are kept as an fp16 staging path plus KVarN records, and the KVarN configuration is stored separately from normal type_k/type_v because each record spans a full 128-token K/V tile.
That is why there are no rows like q8_0-kvarn4 in the current preview. They are not impossible in principle, but they would require a real hybrid-cache architecture: one side allocated and served by the normal KV path, the other side by KVarN, with attention graph routing, CUDA kernels, state save/load, rollback, prompt cache, seq_cp/seq_rm, DFlash backup, SWA/iSWA, and multi-sequence behavior all updated for split ownership. That would be quite complex to implement, and it is unclear whether the partial compression would be worth the extra risk.
Any speed claims also get a caveat. The paper's speed story is about an optimized vLLM-oriented decode path, while the tok/s values in this article are prompt-processing throughput from llama-perplexity. BeeLlama's current path is a raw llama.cpp integration with custom KVarN store/materialize operations and conservative memory handling. If it is slower in prompt processing today, that is operationally true for this build, but it says little about optimized generation speed.
The hardware and model setup matches the original article as closely as possible. Hardware: one RTX 3090 with 24 GB VRAM, Ryzen 7 5700X3D, and 32 GB system RAM. Model: Qwen 3.6 27B. The main comparison is Q5_K_S weights at 64k context, because that is the cleaner weight quant for this purpose. IQ4_XS at 64k and 128k is included mostly as a tie-breaker and noise check; it is useful, but I do not treat it as the primary ranking signal.
The PPL run tested only symmetric KVarN cache types: kvarn4, kvarn3, and kvarn2. The KLD run tested the full 3x3 K/V KVarN matrix: kvarn4-kvarn4, kvarn4-kvarn3, kvarn4-kvarn2, and so on down to kvarn2-kvarn2. KLD also included the control rows from the original benchmark to ensure there's no significant drift from the engine changes.
The follow-up tables at the end use the same three model/context configurations, but only walk the higher KVarN range where the earlier matrix suggested a useful preset could exist: symmetric kvarn5, kvarn6, kvarn8 for PPL, and K≥V pairs for KLD.
All KLD rows are measured against a bf16 KV-cache baseline using llama-perplexity --kl-divergence on Wikitext-2. That is not the paper's pseudo-decode evaluation and it is not an end-to-end reasoning benchmark. It is a local distribution-fidelity test. I like it here because it catches output-distribution movement that PPL tends to hide.
The tok/s column is prompt-processing/prefill throughput from the benchmark run, not generation speed. The table sizes are the actual KV-cache sizes reported by the run logs. KVarN could still behave differently during decode, including faster after optimized kernels; these tables do not decide that question.
PPL makes KVarN look nearly lossless at the top settings. On Q5_K_S 64k, bf16 PPL is 5.4800 and kvarn4 is 5.4807. That looks like nothing happened. But KLD says something did happen: bf16 mean KLD is 0.000375, while kvarn4-kvarn4 is 0.002974. That is not fp16 precision at the distribution level.
That distinction matters. The KVarN paper reports near-fp16 behavior on generative benchmarks such as MATH500, AIME24 and HumanEval. That can be true at the task-score level while still being false at the tensor or distribution level. At 2-4 bits plus stored scales, there is no mathematical reason to expect fp16-equivalent cache reconstruction. What KVarN can do is move the errors into less damaging shapes and preserve the task behavior better than older cache methods.
For this article, the main metric is mean KLD on Q5_K_S 64k, with 99.9% KLD as the tail check and IQ4_XS as a tie-breaker. If a row only wins PPL, I do not count that as much.
Here is the core Q5_K_S 64k comparison. The table is deliberately smaller than the full reference data (available at the end of the article), because this is the part that explains the curve shift.
| Cache | Size vs bf16 | Mean KLD | 99.9% KLD | 99.9% precision vs bf16 | Read |
|---|---|---|---|---|---|
| bf16 | 100.0% | 0.000375 | 0.023258 | 100.00% | Reference |
| q5_1 | 37.5% | 0.002911 | 0.098354 | 92.77% | Best ordinary sub-6-bit row here |
| kvarn4-kvarn4 | 27.9% | 0.002974 | 0.094819 | 93.09% | q5-ish fidelity at q4-ish memory |
| q5_0 | 34.4% | 0.003206 | 0.099073 | 92.70% | Slightly worse mean than KVarN4, larger cache |
| q5_0-q4_0 | 31.3% | 0.003581 | 0.113332 | 91.39% | Old asymmetric middle tier |
| kvarn4-kvarn3 | 24.8% | 0.003824 | 0.135028 | 89.42% | Between q4_0 and q5_0-q4_0, below q4_0 memory |
| q4_0 | 28.1% | 0.004711 | 0.130419 | 89.84% | Old q4 anchor |
| turbo4 | 25.8% | 0.004760 | 0.138370 | 89.13% | Same memory band, worse than kvarn4-kvarn3 |
| kvarn3-kvarn3 | 21.7% | 0.005349 | 0.168135 | 86.51% | Better turbo3-class option |
| turbo3_tcq | 20.3% | 0.007978 | 0.227104 | 81.56% | Old compact recommendation |
| kvarn2-kvarn2 | 15.4% | 0.021395 | 0.630208 | 54.50% | Compression endpoint |
| turbo2_tcq | 14.1% | 0.023073 | 0.632401 | 54.38% | Similar endpoint, slightly worse mean |
kvarn4-kvarn4 is the headline. It almost matches q5_1 on mean KLD, beats q5_0, and uses less memory than q4_0. The q5_1 edge is real but narrow: 0.002911 versus 0.002974. In exchange, q5_1 uses 37.5% of bf16 KV while kvarn4-kvarn4 uses 27.9%. For a VRAM-constrained long-context setup, that is not a practical win for q5_1, it's a quality tie-breaker for people with memory to spare.
kvarn4-kvarn3 is an interesting middle point. It is not simply "smaller q4_0". It sits between q4_0 and q5_0-q4_0: worse than q5_0-q4_0 on the two 64k mean-KLD rows, basically tied on the 128k IQ4_XS row, and consistently better than q4_0. Its memory footprint is lower than both: 24.8% on the 64k runs, versus 28.1% for q4_0 and 31.3% for q5_0-q4_0.
kvarn3-kvarn3 is the compact sweet spot if you were previously looking at turbo3_tcq. On the main row it uses 21.7% instead of 20.3% of bf16 KV, but the mean KLD drops from 0.007978 to 0.005349 and the 99.9% KLD drops from 0.227104 to 0.168135. That is a large quality improvement for a small memory increase.
kvarn2-kvarn2 is not the revolution. It is useful as the low endpoint, and it beats turbo2_tcq on Q5 mean KLD, but the tails are not clean enough to make it the row I would recommend first.
The full 3x3 matrix shows why KVarN should not be treated as a single symmetric setting. The first rule is familiar: K matters more than V. At the same total size, reducing K hurts more than reducing V.
| Q5_K_S 64k pair | Size vs bf16 | Mean KLD | 99.9% KLD | 99.9% precision vs bf16 | Read |
|---|---|---|---|---|---|
| kvarn3-kvarn4 | 24.8% | 0.004652 | 0.140358 | 88.95% | Lower K, higher V |
| kvarn4-kvarn3 | 24.8% | 0.003824 | 0.135028 | 89.42% | Higher K wins |
| kvarn2-kvarn4 | 21.7% | 0.013639 | 0.418240 | 67.37% | 2-bit K is painful |
| kvarn3-kvarn3 | 21.7% | 0.005349 | 0.168135 | 86.51% | Balanced middle wins |
| kvarn4-kvarn2 | 21.7% | 0.010449 | 0.340392 | 72.82% | 2-bit V is also painful |
| kvarn2-kvarn3 | 18.6% | 0.014589 | 0.445014 | 65.59% | Lower K, higher V |
| kvarn3-kvarn2 | 18.6% | 0.011122 | 0.345995 | 72.42% | Higher K wins again |
The second rule is more important: do not overreact and dump all the budget into K. At 21.7% memory, kvarn3-kvarn3 crushes both asymmetric extremes. kvarn2-kvarn4 damages K too much, while kvarn4-kvarn2 damages V too much. K deserves priority, but V cannot be abandoned.
That is why the two settings that matter most are kvarn4-kvarn4 and kvarn4-kvarn3. The first gives the q5-like result. The second keeps K at 4 bits, trims V to 3, and creates a memory tier that ordinary q4/q5 cache quantization did not have.
The turbo comparison is where KVarN looks most like the "TurboQuant that was promised". That is not a formal claim about the TurboQuant paper, but a practical read of this benchmark. The old turbo modes were attractive because they reached memory sizes normal q4/q5 could not. The problem was quality leakage. But KVarN reaches the same broad memory bands with much better KLD.
On the main Q5 row, kvarn4-kvarn3 is smaller than turbo4 and better: 24.8% memory and 0.003824 mean KLD versus 25.8% and 0.004760. On the IQ4_XS rows, the same comparison holds directionally. kvarn3-kvarn3 is slightly larger than turbo3_tcq, but much better: 0.005349 versus 0.007978 mean KLD on Q5, with a much better 99.9% tail.
The 2-bit end is less clean. kvarn2-kvarn2 beats turbo2_tcq on Q5 mean KLD, but it's... bad. The tails are still large, and the IQ4_XS checks are noisy enough that I would treat it as an emergency compression setting, not the main result.
First, this does not prove KVarN has fp16 precision. It obviously does not in KLD. kvarn4-kvarn4 is around the q5 cache tier, not around bf16. That is still strong because it gets there at q4-class memory, but the distribution has moved. If an article says "fp16-level" without specifying task score, not tensor fidelity or KLD, it is being loose with the phrase.
Second, this does not prove the paper's reasoning-benchmark results. The paper is about decode-time accumulation and reports results on generative tasks such as MATH500, AIME24, HumanEval and IFEval. This page uses Wikitext PPL and KL divergence against stored bf16 baseline logits. KLD is a strong local fidelity signal, and it is much more sensitive than PPL, but it is not a replacement for running the actual long-generation tasks.
Third, this does not settle speed. The current BeeLlama implementation is too raw, and these tok/s values are prompt-processing numbers rather than generation numbers. The preview path allocates the right cache shapes and produces meaningful quality data, but it is not the final optimized kernel. I would not use today's implementation neither to reject KVarN nor to claim a speedup.
For VRAM-constrained long context, KVarN looks like a real frontier shift. The old sub-6-bit recommendation was a ladder of compromises: q5 if you could afford it, q4 if you needed memory, turbo3_tcq if you needed even more memory and accepted quality loss. KVarN rearranges that.
In the original 2/3/4-bit sweep, kvarn4-kvarn4 is the cleanest result: q5-class fidelity at q4-class memory. If memory is tighter, kvarn4-kvarn3 creates a useful below-q4 tier, and kvarn3-kvarn3 is the better compact candidate if you were previously looking at turbo3_tcq.
The follow-up changes the higher-end recommendation. The symmetric rows are the clean proof of the tier shift, but the near-balanced asymmetric rows are often the practical value presets. If the cache budget can reach the mid-30% range, kvarn5-kvarn5 is the q6-class proof row and kvarn5-kvarn4 is the tighter value row. If the goal is the q8 fidelity tier without the q8 footprint, kvarn6-kvarn6 is the cleaner row, while kvarn6-kvarn5 is the aggressive smaller variant. kvarn8-kvarn8 mostly belongs as a ceiling check, not the row to reach for first.
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn4 | 1144.00 | 27.9% | 5.4807 | 99.99% | 0.03465 | 97.682% +/- 0.042% | 787.55 | 364.80 |
| kvarn3 | 888.00 | 21.7% | 5.4837 | 99.93% | 0.03466 | 97.069% +/- 0.047% | 790.92 | 363.25 |
| kvarn2 | 632.00 | 15.4% | 5.5499 | 98.74% | 0.03523 | 94.086% +/- 0.065% | 796.36 | 360.77 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn4 | 1144.00 | 27.9% | 5.5198 | 99.95% | 0.03500 | 98.562% +/- 0.033% | 825.75 | 347.93 |
| kvarn3 | 888.00 | 21.7% | 5.5210 | 99.93% | 0.03499 | 97.608% +/- 0.042% | 843.84 | 340.47 |
| kvarn2 | 632.00 | 15.4% | 5.5863 | 98.76% | 0.03554 | 94.218% +/- 0.064% | 846.30 | 339.48 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn4 | 2264.00 | 27.6% | 5.2751 | 99.95% | 0.03271 | 98.625% +/- 0.032% | 631.85 | 445.80 |
| kvarn3 | 1752.00 | 21.4% | 5.2803 | 99.85% | 0.03275 | 97.721% +/- 0.041% | 646.56 | 435.66 |
| kvarn2 | 1240.00 | 15.1% | 5.3382 | 98.77% | 0.03323 | 94.648% +/- 0.062% | 651.61 | 432.28 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn4-kvarn4 | 1144.00 | 27.9% | 0.002974 | 99.74% | 0.000176 | 0.005396 | 0.008555 | 0.023857 | 0.094819 | 93.09% | 17.855265 | 97.682% +/- 0.042% | 760.88 | 377.59 |
| kvarn3-kvarn4 | 1016.00 | 24.8% | 0.004652 | 99.57% | 0.000229 | 0.008663 | 0.013730 | 0.035858 | 0.140358 | 88.95% | 13.967189 | 97.236% +/- 0.045% | 770.52 | 372.87 |
| kvarn4-kvarn3 | 1016.00 | 24.8% | 0.003824 | 99.66% | 0.000203 | 0.007034 | 0.011131 | 0.029529 | 0.135028 | 89.42% | 19.541134 | 97.510% +/- 0.043% | 765.23 | 375.44 |
| kvarn2-kvarn4 | 888.00 | 21.7% | 0.013639 | 98.68% | 0.000353 | 0.029264 | 0.047769 | 0.120707 | 0.418240 | 67.37% | 21.646381 | 95.234% +/- 0.059% | 771.78 | 372.25 |
| kvarn3-kvarn3 | 888.00 | 21.7% | 0.005349 | 99.50% | 0.000237 | 0.010358 | 0.016615 | 0.043053 | 0.168135 | 86.51% | 19.511610 | 97.069% +/- 0.047% | 773.12 | 371.61 |
| kvarn4-kvarn2 | 888.00 | 21.7% | 0.010449 | 99.00% | 0.000425 | 0.019061 | 0.029773 | 0.076875 | 0.340392 | 72.82% | 19.988596 | 96.018% +/- 0.054% | 765.57 | 375.28 |
| kvarn2-kvarn3 | 760.00 | 18.6% | 0.014589 | 98.59% | 0.000346 | 0.031786 | 0.051114 | 0.130933 | 0.445014 | 65.59% | 21.722792 | 95.102% +/- 0.060% | 773.83 | 371.27 |
| kvarn3-kvarn2 | 760.00 | 18.6% | 0.011122 | 98.93% | 0.000289 | 0.022725 | 0.035549 | 0.090061 | 0.345995 | 72.42% | 17.628513 | 95.738% +/- 0.056% | 773.65 | 371.36 |
| kvarn2-kvarn2 | 632.00 | 15.4% | 0.021395 | 97.92% | 0.000429 | 0.046774 | 0.074546 | 0.185504 | 0.630208 | 54.50% | 28.287411 | 94.086% +/- 0.065% | 776.81 | 369.85 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn4-kvarn4 | 1144.00 | 27.9% | 0.001245 | 99.89% | 0.000114 | 0.002240 | 0.003625 | 0.009608 | 0.038723 | 96.60% | 10.492991 | 98.562% +/- 0.033% | 806.72 | 356.13 |
| kvarn3-kvarn4 | 1016.00 | 24.8% | 0.002833 | 99.73% | 0.000190 | 0.005513 | 0.009084 | 0.023476 | 0.085286 | 92.21% | 17.753227 | 97.853% +/- 0.040% | 818.20 | 351.14 |
| kvarn4-kvarn3 | 1016.00 | 24.8% | 0.002290 | 99.78% | 0.000200 | 0.003859 | 0.006310 | 0.017061 | 0.074044 | 93.25% | 20.203844 | 98.152% +/- 0.037% | 811.53 | 354.02 |
| kvarn2-kvarn4 | 888.00 | 21.7% | 0.011961 | 98.82% | 0.000204 | 0.026794 | 0.044280 | 0.114201 | 0.377720 | 68.83% | 11.512545 | 95.618% +/- 0.057% | 821.22 | 349.85 |
| kvarn3-kvarn3 | 888.00 | 21.7% | 0.003708 | 99.64% | 0.000225 | 0.007254 | 0.011907 | 0.032099 | 0.119441 | 89.11% | 21.548656 | 97.608% +/- 0.042% | 822.71 | 349.21 |
| kvarn4-kvarn2 | 888.00 | 21.7% | 0.007576 | 99.25% | 0.000190 | 0.015676 | 0.024607 | 0.062398 | 0.260076 | 77.42% | 13.388462 | 96.398% +/- 0.051% | 813.96 | 352.97 |
| kvarn2-kvarn3 | 760.00 | 18.6% | 0.013017 | 98.72% | 0.000290 | 0.028570 | 0.046846 | 0.121875 | 0.412188 | 66.50% | 21.728636 | 95.438% +/- 0.058% | 824.46 | 348.47 |
| kvarn3-kvarn2 | 760.00 | 18.6% | 0.009046 | 99.11% | 0.000135 | 0.019462 | 0.030744 | 0.079035 | 0.318618 | 73.02% | 8.404756 | 95.983% +/- 0.054% | 824.23 | 348.57 |
| kvarn2-kvarn2 | 632.00 | 15.4% | 0.019693 | 98.06% | 0.000279 | 0.044348 | 0.070586 | 0.180558 | 0.595588 | 55.35% | 12.982862 | 94.218% +/- 0.064% | 824.45 | 348.47 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn4-kvarn4 | 2264.00 | 27.6% | 0.000928 | 99.91% | 0.000011 | 0.002018 | 0.003219 | 0.008277 | 0.030136 | 97.04% | 0.835203 | 98.625% +/- 0.032% | 614.31 | 458.53 |
| kvarn3-kvarn4 | 2008.00 | 24.5% | 0.002185 | 99.78% | 0.000026 | 0.004966 | 0.008147 | 0.021840 | 0.072675 | 93.00% | 1.956919 | 97.963% +/- 0.039% | 624.77 | 450.85 |
| kvarn4-kvarn3 | 2008.00 | 24.5% | 0.001543 | 99.85% | 0.000017 | 0.003377 | 0.005502 | 0.014699 | 0.054317 | 94.72% | 1.113039 | 98.352% +/- 0.035% | 616.44 | 456.95 |
| kvarn2-kvarn4 | 1752.00 | 21.4% | 0.009913 | 99.01% | 0.000082 | 0.023162 | 0.038312 | 0.105779 | 0.350418 | 70.44% | 2.715533 | 95.842% +/- 0.055% | 626.81 | 449.39 |
| kvarn3-kvarn3 | 1752.00 | 21.4% | 0.002776 | 99.72% | 0.000023 | 0.006321 | 0.010481 | 0.028174 | 0.098882 | 90.59% | 0.613806 | 97.721% +/- 0.041% | 627.52 | 448.88 |
| kvarn4-kvarn2 | 1752.00 | 21.4% | 0.006183 | 99.38% | 0.000050 | 0.013974 | 0.021954 | 0.055107 | 0.202573 | 81.67% | 1.954295 | 96.714% +/- 0.049% | 618.51 | 455.42 |
| kvarn2-kvarn3 | 1496.00 | 18.3% | 0.010737 | 98.93% | 0.000090 | 0.025221 | 0.041905 | 0.111761 | 0.368613 | 69.17% | 3.498158 | 95.653% +/- 0.056% | 630.69 | 446.62 |
| kvarn3-kvarn2 | 1496.00 | 18.3% | 0.007559 | 99.25% | 0.000060 | 0.017252 | 0.027371 | 0.068326 | 0.244888 | 78.28% | 2.008955 | 96.293% +/- 0.052% | 630.77 | 446.56 |
| kvarn2-kvarn2 | 1240.00 | 15.1% | 0.016452 | 98.37% | 0.000128 | 0.038350 | 0.061790 | 0.160103 | 0.540234 | 58.26% | 3.977560 | 94.648% +/- 0.062% | 631.58 | 445.99 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| bf16 | 4096.00 | 100.0% | 5.4800 | 100.00% | 0.03465 | 99.647% +/- 0.016% | 851.75 | 326.30 |
| q8_0 | 2176.00 | 53.1% | 5.4774 | 100.05% | 0.03465 | 97.942% +/- 0.039% | 851.57 | 331.27 |
| q6_0 | 1664.00 | 40.6% | 5.4778 | 100.04% | 0.03465 | 97.890% +/- 0.040% | 852.96 | 336.74 |
| q5_1 | 1536.00 | 37.5% | 5.4777 | 100.04% | 0.03464 | 97.787% +/- 0.041% | 848.27 | 332.64 |
| q5_0 | 1408.00 | 34.4% | 5.4802 | 100.00% | 0.03466 | 97.707% +/- 0.041% | 848.36 | 332.45 |
| q4_1 | 1280.00 | 31.3% | 5.4808 | 99.99% | 0.03467 | 97.259% +/- 0.045% | 853.49 | 330.43 |
| q4_0 | 1152.00 | 28.1% | 5.4877 | 99.86% | 0.03473 | 97.179% +/- 0.046% | 853.50 | 330.76 |
| turbo4 | 1056.00 | 25.8% | 5.4841 | 99.93% | 0.03468 | 97.037% +/- 0.047% | 705.06 | 395.16 |
| turbo3_tcq | 832.00 | 20.3% | 5.5054 | 99.54% | 0.03480 | 96.265% +/- 0.052% | 794.21 | 353.31 |
| turbo3 | 800.00 | 19.5% | 5.5149 | 99.37% | 0.03493 | 95.517% +/- 0.057% | 802.71 | 344.83 |
| turbo2_tcq | 576.00 | 14.1% | 5.5705 | 98.38% | 0.03566 | 93.456% +/- 0.068% | 805.25 | 348.78 |
| turbo2 | 544.00 | 13.3% | 5.6403 | 97.16% | 0.03581 | 91.646% +/- 0.076% | 840.34 | 335.07 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| bf16 | 4096.00 | 100.0% | 5.5169 | 100.00% | 0.03497 | 99.776% +/- 0.013% | 909.83 | 336.43 |
| q8_0 | 2176.00 | 53.1% | 5.5157 | 100.02% | 0.03499 | 98.950% +/- 0.028% | 910.43 | 309.49 |
| q6_0 | 1664.00 | 40.6% | 5.5171 | 100.00% | 0.03500 | 98.878% +/- 0.029% | 922.50 | 311.45 |
| q5_1 | 1536.00 | 37.5% | 5.5181 | 99.98% | 0.03501 | 98.618% +/- 0.032% | 906.38 | 310.47 |
| q5_0 | 1408.00 | 34.4% | 5.5175 | 99.99% | 0.03500 | 98.553% +/- 0.033% | 906.88 | 310.27 |
| q4_1 | 1280.00 | 31.3% | 5.5237 | 99.88% | 0.03505 | 97.880% +/- 0.040% | 911.35 | 308.85 |
| q4_0 | 1152.00 | 28.1% | 5.5251 | 99.85% | 0.03505 | 97.793% +/- 0.041% | 912.54 | 308.61 |
| turbo4 | 1056.00 | 25.8% | 5.5277 | 99.80% | 0.03508 | 97.652% +/- 0.042% | 746.00 | 372.98 |
| turbo3_tcq | 832.00 | 20.3% | 5.5426 | 99.54% | 0.03513 | 96.569% +/- 0.050% | 844.33 | 331.64 |
| turbo3 | 800.00 | 19.5% | 5.5561 | 99.29% | 0.03533 | 95.746% +/- 0.056% | 882.76 | 318.29 |
| turbo2_tcq | 576.00 | 14.1% | 5.6085 | 98.37% | 0.03599 | 93.669% +/- 0.067% | 858.66 | 326.24 |
| turbo2 | 544.00 | 13.3% | 5.6823 | 97.09% | 0.03621 | 91.865% +/- 0.076% | 898.66 | 312.75 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| bf16 | 8192.00 | 100.0% | 5.2724 | 100.00% | 0.03269 | 99.995% +/- 0.002% | 703.61 | 389.94 |
| q8_0 | 4352.00 | 53.1% | 5.2716 | 100.02% | 0.03271 | 98.950% +/- 0.028% | 707.02 | 387.84 |
| q6_0 | 3328.00 | 40.6% | 5.2729 | 99.99% | 0.03272 | 98.855% +/- 0.029% | 720.53 | 390.94 |
| q5_1 | 3072.00 | 37.5% | 5.2723 | 100.00% | 0.03271 | 98.603% +/- 0.032% | 702.53 | 390.33 |
| q5_0 | 2816.00 | 34.4% | 5.2738 | 99.97% | 0.03272 | 98.543% +/- 0.033% | 703.18 | 390.05 |
| q4_1 | 2560.00 | 31.3% | 5.2772 | 99.91% | 0.03276 | 97.961% +/- 0.039% | 708.53 | 387.23 |
| q4_0 | 2304.00 | 28.1% | 5.2803 | 99.85% | 0.03276 | 97.793% +/- 0.041% | 709.40 | 386.58 |
| turbo4 | 2112.00 | 25.8% | 5.2822 | 99.81% | 0.03281 | 97.639% +/- 0.042% | 520.46 | 520.82 |
| turbo3_tcq | 1664.00 | 20.3% | 5.2985 | 99.51% | 0.03281 | 96.591% +/- 0.050% | 647.40 | 421.95 |
| turbo3 | 1600.00 | 19.5% | 5.3084 | 99.32% | 0.03301 | 95.861% +/- 0.055% | 689.73 | 396.86 |
| turbo2_tcq | 1152.00 | 14.1% | 5.3513 | 98.53% | 0.03363 | 93.807% +/- 0.067% | 655.44 | 416.90 |
| turbo2 | 1088.00 | 13.3% | 5.4287 | 97.12% | 0.03386 | 92.033% +/- 0.075% | 696.53 | 393.24 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bf16 | 4096.00 | 100.0% | 0.000375 | 100.00% | 0.000058 | 0.000568 | 0.001693 | 0.005234 | 0.023258 | 100.00% | 7.374046 | 99.647% +/- 0.016% | 850.81 | 337.59 |
| bf16-q8_0 | 3136.00 | 76.6% | 0.002475 | 99.79% | 0.000171 | 0.004209 | 0.006629 | 0.019275 | 0.079827 | 94.50% | 14.729765 | 97.991% +/- 0.039% | 850.62 | 337.75 |
| bf16-q6_0 | 2880.00 | 70.3% | 0.002393 | 99.80% | 0.000101 | 0.004302 | 0.006788 | 0.019416 | 0.078853 | 94.59% | 6.564770 | 97.961% +/- 0.039% | 848.99 | 338.40 |
| bf16-q5_1 | 2816.00 | 68.8% | 0.002440 | 99.79% | 0.000090 | 0.004567 | 0.007171 | 0.020056 | 0.096805 | 92.91% | 9.260141 | 97.903% +/- 0.040% | 848.20 | 338.72 |
| bf16-q5_0 | 2752.00 | 67.2% | 0.002491 | 99.79% | 0.000128 | 0.004599 | 0.007110 | 0.020664 | 0.089753 | 93.57% | 13.079553 | 97.829% +/- 0.040% | 848.53 | 338.58 |
| bf16-q4_1 | 2688.00 | 65.6% | 0.003221 | 99.72% | 0.000208 | 0.005661 | 0.008682 | 0.023404 | 0.105923 | 92.07% | 20.044716 | 97.637% +/- 0.042% | 853.30 | 336.69 |
| bf16-q4_0 | 2624.00 | 64.1% | 0.003339 | 99.70% | 0.000187 | 0.005952 | 0.009136 | 0.025430 | 0.110166 | 91.68% | 19.452881 | 97.594% +/- 0.042% | 854.17 | 336.35 |
| bf16-turbo4 | 2576.00 | 62.9% | 0.003549 | 99.68% | 0.000163 | 0.006532 | 0.009849 | 0.026539 | 0.101054 | 92.52% | 11.934117 | 97.473% +/- 0.043% | 843.17 | 340.74 |
| bf16-turbo3_tcq | 2464.00 | 60.2% | 0.005377 | 99.50% | 0.000281 | 0.009815 | 0.014509 | 0.036872 | 0.146175 | 88.43% | 23.307222 | 96.962% +/- 0.047% | 823.83 | 348.74 |
| q8_0 | 2176.00 | 53.1% | 0.002328 | 99.80% | 0.000125 | 0.004233 | 0.006656 | 0.019669 | 0.078709 | 94.61% | 14.355996 | 97.942% +/- 0.039% | 851.11 | 337.66 |
| q8_0-q6_0 | 1920.00 | 46.9% | 0.002499 | 99.79% | 0.000184 | 0.004295 | 0.006708 | 0.019381 | 0.081616 | 94.33% | 17.816196 | 97.942% +/- 0.039% | 848.78 | 338.40 |
| q8_0-q5_1 | 1856.00 | 45.3% | 0.002529 | 99.78% | 0.000143 | 0.004557 | 0.007108 | 0.020346 | 0.082880 | 94.21% | 15.367683 | 97.884% +/- 0.040% | 828.63 | 346.71 |
| q8_0-q5_0 | 1792.00 | 43.8% | 0.002656 | 99.77% | 0.000168 | 0.004673 | 0.007348 | 0.021039 | 0.088486 | 93.69% | 17.987650 | 97.826% +/- 0.040% | 847.33 | 338.90 |
| q8_0-q4_1 | 1728.00 | 42.2% | 0.003080 | 99.73% | 0.000115 | 0.005645 | 0.008587 | 0.023390 | 0.099080 | 92.70% | 8.073231 | 97.655% +/- 0.042% | 786.54 | 364.58 |
| q8_0-q4_0 | 1664.00 | 40.6% | 0.003316 | 99.71% | 0.000165 | 0.005976 | 0.009075 | 0.024892 | 0.104680 | 92.18% | 13.481506 | 97.532% +/- 0.043% | 849.37 | 338.13 |
| q6_0 | 1664.00 | 40.6% | 0.002614 | 99.78% | 0.000180 | 0.004426 | 0.006949 | 0.020078 | 0.090800 | 93.47% | 14.112586 | 97.890% +/- 0.040% | 845.96 | 339.52 |
| q8_0-turbo4 | 1616.00 | 39.5% | 0.003561 | 99.68% | 0.000215 | 0.006518 | 0.009834 | 0.026426 | 0.103041 | 92.33% | 23.102724 | 97.460% +/- 0.043% | 838.90 | 342.38 |
| q6_0-q5_1 | 1600.00 | 39.1% | 0.002781 | 99.76% | 0.000228 | 0.004682 | 0.007348 | 0.020998 | 0.090447 | 93.50% | 23.770491 | 97.913% +/- 0.039% | 846.24 | 339.41 |
| q5_1 | 1536.00 | 37.5% | 0.002911 | 99.75% | 0.000167 | 0.005045 | 0.007916 | 0.022604 | 0.098354 | 92.77% | 13.397068 | 97.787% +/- 0.041% | 841.65 | 341.63 |
| q6_0-q5_0 | 1536.00 | 37.5% | 0.002820 | 99.76% | 0.000209 | 0.004748 | 0.007457 | 0.021883 | 0.092682 | 93.29% | 23.186867 | 97.788% +/- 0.041% | 846.86 | 339.16 |
| q8_0-turbo3_tcq | 1504.00 | 36.7% | 0.005090 | 99.53% | 0.000188 | 0.009736 | 0.014401 | 0.037056 | 0.149387 | 88.15% | 20.128752 | 96.899% +/- 0.048% | 817.57 | 350.23 |
| q6_0-q4_1 | 1472.00 | 35.9% | 0.003312 | 99.71% | 0.000232 | 0.005755 | 0.008847 | 0.024387 | 0.104582 | 92.19% | 23.244659 | 97.605% +/- 0.042% | 848.42 | 338.54 |
| q5_0 | 1408.00 | 34.4% | 0.003206 | 99.72% | 0.000286 | 0.005232 | 0.008194 | 0.022759 | 0.099073 | 92.70% | 22.619892 | 97.707% +/- 0.041% | 849.79 | 338.00 |
| q5_1-q4_1 | 1408.00 | 34.4% | 0.003380 | 99.70% | 0.000195 | 0.006140 | 0.009479 | 0.025886 | 0.095011 | 93.08% | 21.394011 | 97.529% +/- 0.043% | 846.27 | 339.25 |
| q6_0-q4_0 | 1408.00 | 34.4% | 0.003288 | 99.71% | 0.000129 | 0.006096 | 0.009294 | 0.025456 | 0.111566 | 91.55% | 10.711100 | 97.524% +/- 0.043% | 848.24 | 338.61 |
| q6_0-turbo4 | 1360.00 | 33.2% | 0.003748 | 99.66% | 0.000224 | 0.006642 | 0.009997 | 0.026902 | 0.107377 | 91.93% | 16.445103 | 97.465% +/- 0.043% | 837.77 | 342.84 |
| q5_0-q4_1 | 1344.00 | 32.8% | 0.003471 | 99.69% | 0.000206 | 0.006310 | 0.009582 | 0.025829 | 0.099618 | 92.65% | 21.863117 | 97.539% +/- 0.043% | 847.59 | 339.65 |
| q5_1-q4_0 | 1344.00 | 32.8% | 0.003626 | 99.68% | 0.000212 | 0.006441 | 0.009773 | 0.025668 | 0.108649 | 91.82% | 15.809726 | 97.515% +/- 0.043% | 846.91 | 339.23 |
| q4_1 | 1280.00 | 31.3% | 0.004476 | 99.59% | 0.000267 | 0.007716 | 0.011901 | 0.031166 | 0.141813 | 88.82% | 18.150869 | 97.259% +/- 0.045% | 854.33 | 336.49 |
| q5_0-q4_0 | 1280.00 | 31.3% | 0.003581 | 99.68% | 0.000174 | 0.006600 | 0.010058 | 0.027423 | 0.113332 | 91.39% | 14.938599 | 97.437% +/- 0.044% | 847.64 | 338.79 |
| q6_0-turbo3_tcq | 1248.00 | 30.5% | 0.005379 | 99.50% | 0.000247 | 0.009906 | 0.014556 | 0.037285 | 0.154680 | 87.68% | 19.739548 | 96.922% +/- 0.048% | 819.23 | 350.60 |
| q5_0-turbo4 | 1232.00 | 30.1% | 0.003812 | 99.66% | 0.000176 | 0.007068 | 0.010735 | 0.028203 | 0.112249 | 91.49% | 17.032024 | 97.371% +/- 0.044% | 837.52 | 342.95 |
| q5_1-turbo3_tcq | 1184.00 | 28.9% | 0.005594 | 99.48% | 0.000291 | 0.010264 | 0.015324 | 0.038175 | 0.144591 | 88.57% | 24.684429 | 96.878% +/- 0.048% | 816.05 | 350.73 |
| q4_0 | 1152.00 | 28.1% | 0.004711 | 99.57% | 0.000301 | 0.008439 | 0.012949 | 0.033663 | 0.130419 | 89.84% | 21.636135 | 97.179% +/- 0.046% | 855.08 | 336.11 |
| q5_0-turbo3_tcq | 1120.00 | 27.3% | 0.005471 | 99.49% | 0.000265 | 0.010259 | 0.015229 | 0.038214 | 0.158514 | 87.35% | 22.268801 | 96.865% +/- 0.048% | 815.80 | 350.94 |
| q5_0-turbo3 | 1104.00 | 27.0% | 0.007097 | 99.33% | 0.000259 | 0.013747 | 0.020259 | 0.048761 | 0.192428 | 84.44% | 18.094296 | 96.331% +/- 0.052% | 837.90 | 342.47 |
| q4_1-turbo3_tcq | 1056.00 | 25.8% | 0.006184 | 99.42% | 0.000292 | 0.011652 | 0.017320 | 0.042997 | 0.174831 | 85.94% | 25.079035 | 96.663% +/- 0.050% | 816.95 | 350.43 |
| turbo4 | 1056.00 | 25.8% | 0.004760 | 99.55% | 0.000201 | 0.009046 | 0.013692 | 0.035205 | 0.138370 | 89.13% | 13.967494 | 97.037% +/- 0.047% | 705.32 | 401.18 |
| q4_0-turbo3_tcq | 992.00 | 24.2% | 0.006269 | 99.41% | 0.000270 | 0.012220 | 0.018173 | 0.045421 | 0.186572 | 84.93% | 23.157375 | 96.622% +/- 0.050% | 821.89 | 349.67 |
| q4_0-turbo3 | 976.00 | 23.8% | 0.008235 | 99.22% | 0.000336 | 0.015576 | 0.022828 | 0.056527 | 0.222154 | 81.96% | 24.353268 | 96.075% +/- 0.054% | 839.29 | 341.78 |
| q4_0-turbo2_tcq | 864.00 | 21.1% | 0.015168 | 98.53% | 0.000288 | 0.031826 | 0.045569 | 0.105461 | 0.395244 | 68.94% | 20.743238 | 94.591% +/- 0.062% | 826.07 | 347.04 |
| turbo3_tcq | 832.00 | 20.3% | 0.007978 | 99.24% | 0.000267 | 0.015663 | 0.023628 | 0.058286 | 0.227104 | 81.56% | 20.517471 | 96.265% +/- 0.052% | 795.20 | 359.09 |
| turbo3 | 800.00 | 19.5% | 0.011181 | 98.93% | 0.000304 | 0.022805 | 0.034209 | 0.082015 | 0.296060 | 76.12% | 22.977211 | 95.517% +/- 0.057% | 836.75 | 342.73 |
| turbo3_tcq-turbo2_tcq | 704.00 | 17.2% | 0.016386 | 98.41% | 0.000283 | 0.034186 | 0.049133 | 0.115072 | 0.437043 | 66.11% | 18.275532 | 94.379% +/- 0.064% | 796.16 | 358.86 |
| turbo3-turbo2 | 672.00 | 16.4% | 0.023985 | 97.67% | 0.000403 | 0.050100 | 0.072850 | 0.168258 | 0.605087 | 55.89% | 20.812553 | 93.154% +/- 0.070% | 831.88 | 344.85 |
| turbo2_tcq | 576.00 | 14.1% | 0.023073 | 97.76% | 0.000420 | 0.048777 | 0.071865 | 0.170350 | 0.632401 | 54.38% | 24.771320 | 93.456% +/- 0.068% | 807.25 | 354.12 |
| turbo2 | 544.00 | 13.3% | 0.036230 | 96.48% | 0.000465 | 0.078942 | 0.117545 | 0.276438 | 0.903576 | 41.47% | 26.508263 | 91.646% +/- 0.076% | 842.29 | 340.66 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bf16 | 4096.00 | 100.0% | 0.000097 | 100.00% | 0.000020 | 0.000186 | 0.000398 | 0.001062 | 0.004152 | 100.00% | 2.345056 | 99.776% +/- 0.013% | 909.80 | 315.80 |
| bf16-q8_0 | 3136.00 | 76.6% | 0.000587 | 99.95% | 0.000092 | 0.000924 | 0.001408 | 0.004018 | 0.016148 | 98.81% | 11.615892 | 98.952% +/- 0.028% | 909.80 | 315.78 |
| bf16-q6_0 | 2880.00 | 70.3% | 0.000636 | 99.95% | 0.000079 | 0.001029 | 0.001556 | 0.004148 | 0.017005 | 98.72% | 8.440266 | 98.875% +/- 0.029% | 908.22 | 316.33 |
| bf16-q5_1 | 2816.00 | 68.8% | 0.000757 | 99.93% | 0.000074 | 0.001283 | 0.001921 | 0.005045 | 0.020206 | 98.41% | 8.462649 | 98.798% +/- 0.030% | 907.39 | 316.62 |
| bf16-q5_0 | 2752.00 | 67.2% | 0.000873 | 99.92% | 0.000101 | 0.001384 | 0.002088 | 0.005688 | 0.022394 | 98.19% | 10.605267 | 98.764% +/- 0.031% | 908.46 | 316.25 |
| bf16-q4_1 | 2688.00 | 65.6% | 0.001357 | 99.87% | 0.000095 | 0.002455 | 0.003607 | 0.008977 | 0.034741 | 96.99% | 10.624268 | 98.392% +/- 0.035% | 913.53 | 314.50 |
| bf16-q4_0 | 2624.00 | 64.1% | 0.001459 | 99.86% | 0.000073 | 0.002756 | 0.004021 | 0.010121 | 0.039676 | 96.51% | 6.619313 | 98.367% +/- 0.035% | 914.07 | 314.31 |
| bf16-turbo4 | 2576.00 | 62.9% | 0.001770 | 99.83% | 0.000097 | 0.003286 | 0.004741 | 0.011679 | 0.043127 | 96.18% | 8.782248 | 98.144% +/- 0.037% | 901.27 | 318.77 |
| bf16-turbo3_tcq | 2464.00 | 60.2% | 0.003256 | 99.68% | 0.000092 | 0.006574 | 0.009463 | 0.022278 | 0.084448 | 92.28% | 8.791702 | 97.398% +/- 0.044% | 879.05 | 326.83 |
| q8_0 | 2176.00 | 53.1% | 0.000577 | 99.95% | 0.000073 | 0.000933 | 0.001428 | 0.004050 | 0.017372 | 98.69% | 8.130807 | 98.950% +/- 0.028% | 912.71 | 314.76 |
| q8_0-q6_0 | 1920.00 | 46.9% | 0.000659 | 99.94% | 0.000093 | 0.001032 | 0.001578 | 0.004419 | 0.018670 | 98.56% | 10.625672 | 98.906% +/- 0.029% | 908.70 | 316.18 |
| q8_0-q5_1 | 1856.00 | 45.3% | 0.000836 | 99.93% | 0.000110 | 0.001291 | 0.001929 | 0.005272 | 0.021544 | 98.28% | 11.861989 | 98.814% +/- 0.030% | 895.23 | 320.91 |
| q8_0-q5_0 | 1792.00 | 43.8% | 0.000881 | 99.92% | 0.000107 | 0.001392 | 0.002092 | 0.005574 | 0.022435 | 98.19% | 10.867614 | 98.714% +/- 0.031% | 906.00 | 316.81 |
| q8_0-q4_1 | 1728.00 | 42.2% | 0.001317 | 99.88% | 0.000122 | 0.002436 | 0.003572 | 0.008974 | 0.034706 | 96.99% | 6.675264 | 98.357% +/- 0.035% | 818.78 | 346.30 |
| q8_0-q4_0 | 1664.00 | 40.6% | 0.001606 | 99.85% | 0.000118 | 0.002793 | 0.004084 | 0.009969 | 0.039299 | 96.55% | 8.993986 | 98.309% +/- 0.036% | 908.09 | 316.08 |
| q6_0 | 1664.00 | 40.6% | 0.000766 | 99.93% | 0.000109 | 0.001179 | 0.001791 | 0.004762 | 0.020407 | 98.39% | 11.368995 | 98.878% +/- 0.029% | 906.47 | 316.96 |
| q8_0-turbo4 | 1616.00 | 39.5% | 0.001845 | 99.83% | 0.000108 | 0.003311 | 0.004787 | 0.011551 | 0.046124 | 95.89% | 8.488309 | 98.147% +/- 0.037% | 898.63 | 319.72 |
| q6_0-q5_1 | 1600.00 | 39.1% | 0.000882 | 99.92% | 0.000099 | 0.001431 | 0.002169 | 0.005746 | 0.021968 | 98.23% | 10.728148 | 98.772% +/- 0.030% | 906.67 | 316.89 |
| q5_1 | 1536.00 | 37.5% | 0.001019 | 99.91% | 0.000075 | 0.001787 | 0.002724 | 0.007262 | 0.029854 | 97.46% | 6.707493 | 98.618% +/- 0.032% | 907.45 | 316.44 |
| q6_0-q5_0 | 1536.00 | 37.5% | 0.000933 | 99.92% | 0.000103 | 0.001519 | 0.002269 | 0.006044 | 0.023588 | 98.08% | 10.475766 | 98.666% +/- 0.032% | 906.68 | 316.89 |
| q8_0-turbo3_tcq | 1504.00 | 36.7% | 0.003336 | 99.68% | 0.000119 | 0.006580 | 0.009451 | 0.022374 | 0.084818 | 92.25% | 11.130499 | 97.411% +/- 0.044% | 871.88 | 328.10 |
| q6_0-q4_1 | 1472.00 | 35.9% | 0.001488 | 99.87% | 0.000115 | 0.002593 | 0.003830 | 0.009763 | 0.037581 | 96.71% | 10.889835 | 98.378% +/- 0.035% | 908.06 | 316.40 |
| q5_0 | 1408.00 | 34.4% | 0.001135 | 99.90% | 0.000088 | 0.002028 | 0.003113 | 0.008113 | 0.031348 | 97.32% | 8.001913 | 98.553% +/- 0.033% | 908.72 | 315.93 |
| q5_1-q4_1 | 1408.00 | 34.4% | 0.001683 | 99.84% | 0.000140 | 0.002928 | 0.004329 | 0.010956 | 0.038976 | 96.58% | 11.203689 | 98.302% +/- 0.036% | 906.39 | 316.62 |
| q6_0-q4_0 | 1408.00 | 34.4% | 0.001555 | 99.85% | 0.000113 | 0.002893 | 0.004200 | 0.010392 | 0.039601 | 96.52% | 11.934310 | 98.279% +/- 0.036% | 909.10 | 316.04 |
| q6_0-turbo4 | 1360.00 | 33.2% | 0.001933 | 99.82% | 0.000118 | 0.003445 | 0.005001 | 0.012021 | 0.044722 | 96.02% | 8.842097 | 98.076% +/- 0.038% | 896.11 | 320.62 |
| q5_0-q4_1 | 1344.00 | 32.8% | 0.001529 | 99.86% | 0.000037 | 0.003073 | 0.004593 | 0.011795 | 0.042828 | 96.21% | 2.328933 | 98.227% +/- 0.036% | 905.93 | 316.98 |
| q5_1-q4_0 | 1344.00 | 32.8% | 0.001813 | 99.83% | 0.000163 | 0.003236 | 0.004782 | 0.011742 | 0.042893 | 96.20% | 18.077213 | 98.160% +/- 0.037% | 905.51 | 317.08 |
| q4_1 | 1280.00 | 31.3% | 0.002316 | 99.78% | 0.000104 | 0.004441 | 0.006776 | 0.016734 | 0.068858 | 93.73% | 8.874204 | 97.880% +/- 0.040% | 913.32 | 314.50 |
| q5_0-q4_0 | 1280.00 | 31.3% | 0.001936 | 99.82% | 0.000147 | 0.003368 | 0.005013 | 0.012419 | 0.044393 | 96.06% | 14.364779 | 98.125% +/- 0.037% | 906.57 | 316.61 |
| q6_0-turbo3_tcq | 1248.00 | 30.5% | 0.003412 | 99.67% | 0.000121 | 0.006702 | 0.009642 | 0.022886 | 0.089874 | 91.78% | 9.695003 | 97.394% +/- 0.044% | 874.99 | 328.36 |
| q5_0-turbo4 | 1232.00 | 30.1% | 0.002122 | 99.80% | 0.000131 | 0.003913 | 0.005769 | 0.013995 | 0.052315 | 95.30% | 11.286289 | 97.977% +/- 0.039% | 895.90 | 320.70 |
| q5_1-turbo3_tcq | 1184.00 | 28.9% | 0.003560 | 99.65% | 0.000130 | 0.007025 | 0.010176 | 0.024077 | 0.088706 | 91.89% | 10.154081 | 97.304% +/- 0.045% | 870.78 | 328.52 |
| q4_0 | 1152.00 | 28.1% | 0.002759 | 99.73% | 0.000141 | 0.005219 | 0.007950 | 0.019862 | 0.076663 | 93.01% | 10.045764 | 97.793% +/- 0.041% | 914.36 | 314.20 |
| q5_0-turbo3_tcq | 1120.00 | 27.3% | 0.003600 | 99.65% | 0.000099 | 0.007198 | 0.010485 | 0.024752 | 0.102109 | 90.67% | 9.033820 | 97.295% +/- 0.045% | 869.81 | 328.94 |
| q5_0-turbo3 | 1104.00 | 27.0% | 0.005209 | 99.49% | 0.000150 | 0.010602 | 0.015245 | 0.036580 | 0.134359 | 87.79% | 11.506174 | 96.750% +/- 0.049% | 894.86 | 320.42 |
| q4_1-turbo3_tcq | 1056.00 | 25.8% | 0.004226 | 99.59% | 0.000107 | 0.008465 | 0.012513 | 0.029618 | 0.121854 | 88.90% | 8.723818 | 97.117% +/- 0.046% | 871.93 | 328.15 |
| turbo4 | 1056.00 | 25.8% | 0.002988 | 99.71% | 0.000130 | 0.005881 | 0.008839 | 0.021868 | 0.076363 | 93.03% | 9.168183 | 97.652% +/- 0.042% | 744.85 | 379.35 |
| q4_0-turbo3_tcq | 992.00 | 24.2% | 0.004466 | 99.56% | 0.000123 | 0.009097 | 0.013470 | 0.032288 | 0.108662 | 90.08% | 9.383696 | 97.067% +/- 0.047% | 871.33 | 328.57 |
| q4_0-turbo3 | 976.00 | 23.8% | 0.006007 | 99.41% | 0.000136 | 0.012352 | 0.018044 | 0.042428 | 0.161644 | 85.43% | 11.092446 | 96.508% +/- 0.051% | 897.34 | 319.56 |
| q4_0-turbo2_tcq | 864.00 | 21.1% | 0.013595 | 98.66% | 0.000204 | 0.028687 | 0.041321 | 0.095393 | 0.367825 | 69.51% | 14.222522 | 94.742% +/- 0.062% | 881.14 | 325.07 |
| turbo3_tcq | 832.00 | 20.3% | 0.006038 | 99.41% | 0.000127 | 0.012714 | 0.019058 | 0.046219 | 0.172480 | 84.51% | 9.415985 | 96.569% +/- 0.050% | 845.55 | 337.45 |
| turbo3 | 800.00 | 19.5% | 0.009102 | 99.10% | 0.000164 | 0.019502 | 0.029311 | 0.068266 | 0.236472 | 79.27% | 10.847077 | 95.746% +/- 0.056% | 894.11 | 320.59 |
| turbo3_tcq-turbo2_tcq | 704.00 | 17.2% | 0.014461 | 98.57% | 0.000165 | 0.031014 | 0.045330 | 0.104559 | 0.374854 | 69.02% | 10.150832 | 94.578% +/- 0.063% | 847.59 | 336.76 |
| turbo3-turbo2 | 672.00 | 16.4% | 0.022168 | 97.82% | 0.000271 | 0.046698 | 0.068008 | 0.160327 | 0.602649 | 54.96% | 18.985191 | 93.434% +/- 0.068% | 884.98 | 323.44 |
| turbo2_tcq | 576.00 | 14.1% | 0.020739 | 97.96% | 0.000230 | 0.045497 | 0.068026 | 0.161256 | 0.538190 | 58.62% | 15.582352 | 93.669% +/- 0.067% | 861.17 | 331.75 |
| turbo2 | 544.00 | 13.3% | 0.034380 | 96.63% | 0.000340 | 0.075876 | 0.113734 | 0.265535 | 0.895385 | 41.01% | 19.482079 | 91.865% +/- 0.076% | 901.01 | 318.44 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | KLD +/- | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| bf16 | 8192.00 | 100.0% | 0.000000 | 100.00% | 0.000000 | 0.000015 | 0.000023 | 0.000037 | 0.000051 | 100.00% | 0.000067 | 99.995% +/- 0.002% | 702.50 | 400.97 |
| bf16-q8_0 | 6272.00 | 76.6% | 0.000479 | 99.95% | 0.000011 | 0.000971 | 0.001473 | 0.003985 | 0.014957 | 98.52% | 1.255583 | 98.953% +/- 0.028% | 705.73 | 399.13 |
| bf16-q6_0 | 5760.00 | 70.3% | 0.000528 | 99.95% | 0.000011 | 0.001080 | 0.001640 | 0.004261 | 0.016915 | 98.33% | 1.246464 | 98.920% +/- 0.029% | 703.90 | 400.17 |
| bf16-q5_1 | 5632.00 | 68.8% | 0.000643 | 99.94% | 0.000007 | 0.001318 | 0.001981 | 0.005192 | 0.019604 | 98.06% | 0.459653 | 98.803% +/- 0.030% | 703.75 | 400.25 |
| bf16-q5_0 | 5504.00 | 67.2% | 0.000680 | 99.93% | 0.000007 | 0.001410 | 0.002094 | 0.005333 | 0.020995 | 97.93% | 0.448673 | 98.758% +/- 0.031% | 703.91 | 400.16 |
| bf16-q4_1 | 5376.00 | 65.6% | 0.001145 | 99.89% | 0.000011 | 0.002439 | 0.003558 | 0.008389 | 0.030617 | 96.99% | 0.856124 | 98.425% +/- 0.034% | 709.30 | 397.13 |
| bf16-q4_0 | 5248.00 | 64.1% | 0.001295 | 99.87% | 0.000016 | 0.002773 | 0.004027 | 0.009434 | 0.034897 | 96.58% | 1.481670 | 98.347% +/- 0.035% | 710.30 | 396.57 |
| bf16-turbo4 | 5152.00 | 62.9% | 0.001533 | 99.85% | 0.000013 | 0.003304 | 0.004779 | 0.011221 | 0.041510 | 95.94% | 0.505207 | 98.119% +/- 0.038% | 697.20 | 404.02 |
| bf16-turbo3_tcq | 4928.00 | 60.2% | 0.003191 | 99.68% | 0.000027 | 0.006856 | 0.009858 | 0.023102 | 0.086196 | 91.75% | 1.071732 | 97.420% +/- 0.044% | 678.44 | 415.19 |
| q8_0 | 4352.00 | 53.1% | 0.000482 | 99.95% | 0.000007 | 0.000983 | 0.001508 | 0.004061 | 0.014951 | 98.52% | 0.478175 | 98.950% +/- 0.028% | 708.31 | 397.81 |
| q8_0-q6_0 | 3840.00 | 46.9% | 0.000530 | 99.95% | 0.000006 | 0.001093 | 0.001652 | 0.004335 | 0.017408 | 98.28% | 0.468998 | 98.900% +/- 0.029% | 703.96 | 400.14 |
| q8_0-q5_1 | 3712.00 | 45.3% | 0.000651 | 99.93% | 0.000012 | 0.001335 | 0.002010 | 0.005161 | 0.018918 | 98.13% | 1.270212 | 98.779% +/- 0.030% | 694.87 | 405.20 |
| q8_0-q5_0 | 3584.00 | 43.8% | 0.000703 | 99.93% | 0.000013 | 0.001433 | 0.002141 | 0.005523 | 0.020360 | 97.99% | 1.166986 | 98.757% +/- 0.031% | 702.31 | 400.84 |
| q8_0-q4_1 | 3456.00 | 42.2% | 0.001149 | 99.89% | 0.000012 | 0.002453 | 0.003582 | 0.008568 | 0.029970 | 97.05% | 0.964733 | 98.407% +/- 0.035% | 637.52 | 440.27 |
| q8_0-q4_0 | 3328.00 | 40.6% | 0.001295 | 99.87% | 0.000016 | 0.002765 | 0.003998 | 0.009587 | 0.035741 | 96.49% | 1.614931 | 98.304% +/- 0.036% | 706.17 | 398.89 |
| q6_0 | 3328.00 | 40.6% | 0.000589 | 99.94% | 0.000008 | 0.001212 | 0.001852 | 0.004706 | 0.019175 | 98.11% | 0.573200 | 98.855% +/- 0.029% | 701.36 | 401.62 |
| q8_0-turbo4 | 3232.00 | 39.5% | 0.001554 | 99.84% | 0.000014 | 0.003335 | 0.004810 | 0.011473 | 0.041006 | 95.99% | 0.633252 | 98.171% +/- 0.037% | 694.04 | 405.86 |
| q6_0-q5_1 | 3200.00 | 39.1% | 0.000703 | 99.93% | 0.000011 | 0.001452 | 0.002180 | 0.005544 | 0.021821 | 97.85% | 1.145905 | 98.740% +/- 0.031% | 701.05 | 401.80 |
| q5_1 | 3072.00 | 37.5% | 0.000827 | 99.92% | 0.000008 | 0.001792 | 0.002687 | 0.006764 | 0.023291 | 97.70% | 0.496846 | 98.603% +/- 0.032% | 702.81 | 400.71 |
| q6_0-q5_0 | 3072.00 | 37.5% | 0.000752 | 99.92% | 0.000013 | 0.001552 | 0.002311 | 0.005872 | 0.022227 | 97.81% | 1.083445 | 98.700% +/- 0.031% | 701.69 | 401.43 |
| q8_0-turbo3_tcq | 3008.00 | 36.7% | 0.003167 | 99.68% | 0.000029 | 0.006791 | 0.009860 | 0.022935 | 0.081350 | 92.19% | 2.329764 | 97.407% +/- 0.044% | 672.90 | 417.21 |
| q6_0-q4_1 | 2944.00 | 35.9% | 0.001191 | 99.88% | 0.000009 | 0.002583 | 0.003791 | 0.008875 | 0.031032 | 96.95% | 0.431630 | 98.412% +/- 0.035% | 703.37 | 400.48 |
| q5_0 | 2816.00 | 34.4% | 0.000926 | 99.91% | 0.000007 | 0.002037 | 0.003067 | 0.007468 | 0.027410 | 97.30% | 0.427949 | 98.543% +/- 0.033% | 704.01 | 400.14 |
| q5_1-q4_1 | 2816.00 | 34.4% | 0.001335 | 99.87% | 0.000013 | 0.002884 | 0.004221 | 0.010047 | 0.035062 | 96.56% | 0.850439 | 98.334% +/- 0.035% | 703.75 | 400.05 |
| q6_0-q4_0 | 2816.00 | 34.4% | 0.001317 | 99.87% | 0.000010 | 0.002875 | 0.004183 | 0.009784 | 0.031863 | 96.87% | 0.504899 | 98.275% +/- 0.036% | 704.62 | 399.76 |
| q6_0-turbo4 | 2720.00 | 33.2% | 0.001590 | 99.84% | 0.000016 | 0.003435 | 0.004979 | 0.011859 | 0.039001 | 96.18% | 1.399161 | 98.063% +/- 0.038% | 691.93 | 407.09 |
| q5_0-q4_1 | 2688.00 | 32.8% | 0.001387 | 99.86% | 0.000013 | 0.003017 | 0.004437 | 0.010411 | 0.035706 | 96.50% | 0.971476 | 98.243% +/- 0.036% | 702.90 | 400.58 |
| q5_1-q4_0 | 2688.00 | 32.8% | 0.001485 | 99.85% | 0.000014 | 0.003200 | 0.004746 | 0.011342 | 0.040530 | 96.03% | 0.976738 | 98.200% +/- 0.037% | 702.58 | 400.94 |
| q4_1 | 2560.00 | 31.3% | 0.001933 | 99.81% | 0.000013 | 0.004318 | 0.006595 | 0.016046 | 0.050918 | 95.04% | 0.435837 | 97.961% +/- 0.039% | 709.38 | 397.11 |
| q5_0-q4_0 | 2560.00 | 31.3% | 0.001529 | 99.85% | 0.000016 | 0.003316 | 0.004868 | 0.011640 | 0.039033 | 96.18% | 1.116606 | 98.211% +/- 0.037% | 704.22 | 399.88 |
| q6_0-turbo3_tcq | 2496.00 | 30.5% | 0.003238 | 99.68% | 0.000031 | 0.006922 | 0.009933 | 0.023271 | 0.087341 | 91.64% | 2.346832 | 97.388% +/- 0.044% | 673.13 | 418.46 |
| q5_0-turbo4 | 2464.00 | 30.1% | 0.001809 | 99.82% | 0.000023 | 0.003891 | 0.005742 | 0.013425 | 0.047328 | 95.38% | 2.025323 | 98.004% +/- 0.039% | 691.62 | 407.28 |
| q5_1-turbo3_tcq | 2368.00 | 28.9% | 0.003360 | 99.66% | 0.000029 | 0.007229 | 0.010496 | 0.025104 | 0.089474 | 91.45% | 2.237170 | 97.375% +/- 0.044% | 670.63 | 418.53 |
| q4_0 | 2304.00 | 28.1% | 0.002259 | 99.77% | 0.000017 | 0.005058 | 0.007697 | 0.018505 | 0.058301 | 94.34% | 1.074671 | 97.793% +/- 0.041% | 710.51 | 396.57 |
| q5_0-turbo3_tcq | 2240.00 | 27.3% | 0.003391 | 99.66% | 0.000030 | 0.007321 | 0.010567 | 0.024422 | 0.090901 | 91.32% | 2.252987 | 97.384% +/- 0.044% | 670.54 | 418.68 |
| q5_0-turbo3 | 2208.00 | 27.0% | 0.004728 | 99.53% | 0.000035 | 0.010375 | 0.014767 | 0.034198 | 0.121340 | 88.58% | 1.809964 | 96.732% +/- 0.049% | 693.49 | 405.71 |
| q4_1-turbo3_tcq | 2112.00 | 25.8% | 0.003981 | 99.60% | 0.000034 | 0.008612 | 0.012701 | 0.030111 | 0.112812 | 89.34% | 2.193686 | 97.182% +/- 0.046% | 672.84 | 417.30 |
| turbo4 | 2112.00 | 25.8% | 0.002605 | 99.74% | 0.000024 | 0.005701 | 0.008653 | 0.020367 | 0.071902 | 93.07% | 1.179263 | 97.639% +/- 0.042% | 519.76 | 531.96 |
| q4_0-turbo3_tcq | 1984.00 | 24.2% | 0.004131 | 99.59% | 0.000032 | 0.009062 | 0.013401 | 0.031435 | 0.112297 | 89.38% | 2.139016 | 97.078% +/- 0.047% | 671.63 | 418.13 |
| q4_0-turbo3 | 1952.00 | 23.8% | 0.005488 | 99.45% | 0.000040 | 0.012073 | 0.017644 | 0.040047 | 0.145158 | 86.49% | 1.545795 | 96.549% +/- 0.050% | 695.78 | 404.37 |
| q4_0-turbo2_tcq | 1728.00 | 21.1% | 0.013329 | 98.68% | 0.000090 | 0.029805 | 0.042793 | 0.096811 | 0.306630 | 73.60% | 6.779726 | 94.713% +/- 0.062% | 678.10 | 414.30 |
| turbo3_tcq | 1664.00 | 20.3% | 0.005708 | 99.43% | 0.000045 | 0.012748 | 0.019441 | 0.045320 | 0.150144 | 86.06% | 2.079510 | 96.591% +/- 0.050% | 647.62 | 432.43 |
| turbo3 | 1600.00 | 19.5% | 0.008334 | 99.17% | 0.000057 | 0.019023 | 0.028596 | 0.066157 | 0.207468 | 81.27% | 2.454834 | 95.861% +/- 0.055% | 691.41 | 406.63 |
| turbo3_tcq-turbo2_tcq | 1408.00 | 17.2% | 0.014344 | 98.58% | 0.000086 | 0.032243 | 0.046737 | 0.105269 | 0.343951 | 70.90% | 4.010866 | 94.530% +/- 0.063% | 648.06 | 432.26 |
| turbo3-turbo2 | 1344.00 | 16.4% | 0.020468 | 97.97% | 0.000154 | 0.045514 | 0.065976 | 0.147745 | 0.474415 | 62.23% | 11.938387 | 93.540% +/- 0.068% | 686.57 | 409.17 |
| turbo2_tcq | 1152.00 | 14.1% | 0.019857 | 98.03% | 0.000122 | 0.045491 | 0.068398 | 0.158766 | 0.453761 | 63.53% | 4.370085 | 93.807% +/- 0.067% | 656.84 | 426.67 |
| turbo2 | 1088.00 | 13.3% | 0.032631 | 96.79% | 0.000203 | 0.073833 | 0.111765 | 0.261607 | 0.838113 | 43.25% | 4.735642 | 92.033% +/- 0.075% | 698.21 | 402.96 |
This follow-up was run on updated BeeLlama v0.3.2 Preview that extends the KVarN matrix above the original 2/3/4-bit sweep. The same three model/context configurations are used: Q5_K_S 64k, IQ4_XS 64k, and IQ4_XS 128k, each measured against its bf16 KV-cache baseline.
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn8 | 2168.00 | 52.9% | 5.4784 | 100.03% | 0.03463 | 97.979% +/- 0.039% | 647.73 | 404.71 |
| kvarn6 | 1656.00 | 40.4% | 5.4779 | 100.04% | 0.03462 | 97.959% +/- 0.039% | 704.44 | 372.13 |
| kvarn5 | 1400.00 | 34.2% | 5.4782 | 100.03% | 0.03462 | 97.912% +/- 0.039% | 729.05 | 359.57 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn8 | 2168.00 | 52.9% | 5.5173 | 99.99% | 0.03498 | 99.038% +/- 0.027% | 685.13 | 382.62 |
| kvarn6 | 1656.00 | 40.4% | 5.5171 | 100.00% | 0.03498 | 98.999% +/- 0.027% | 750.05 | 349.50 |
| kvarn5 | 1400.00 | 34.2% | 5.5182 | 99.98% | 0.03498 | 98.899% +/- 0.029% | 761.54 | 344.23 |
| Cache | KV cache (MiB) | Size vs bf16 | Median PPL | Precision vs bf16 | PPL +/- | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|
| kvarn8 | 4312.00 | 52.6% | 5.2738 | 99.97% | 0.03270 | 99.064% +/- 0.027% | 513.70 | 510.31 |
| kvarn6 | 3288.00 | 40.1% | 5.2733 | 99.98% | 0.03270 | 99.051% +/- 0.027% | 583.14 | 449.54 |
| kvarn5 | 2776.00 | 33.9% | 5.2745 | 99.96% | 0.03271 | 98.907% +/- 0.029% | 593.95 | 441.36 |
These PPL rows are only a sanity check. As in the main article, PPL barely moves across the useful cache range, so the recommendation should come from the KLD tables below rather than from the median PPL column.
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn8-kvarn8 | 2168.00 | 52.9% | 0.002361 | 99.80% | 0.004069 | 0.006427 | 0.018345 | 0.076809 | 94.79% | 17.505587 | 97.979% +/- 0.039% | 634.12 | 413.4 |
| kvarn8-kvarn6 | 1912.00 | 46.7% | 0.002390 | 99.80% | 0.004109 | 0.006465 | 0.019044 | 0.082415 | 94.26% | 17.531494 | 98.003% +/- 0.039% | 643.46 | 407.4 |
| kvarn8-kvarn5 | 1784.00 | 43.6% | 0.002266 | 99.81% | 0.004149 | 0.006553 | 0.018835 | 0.084573 | 94.05% | 10.895293 | 98.006% +/- 0.039% | 646.63 | 405.4 |
| kvarn6-kvarn6 | 1656.00 | 40.4% | 0.002338 | 99.80% | 0.004134 | 0.006486 | 0.019045 | 0.078797 | 94.6% | 13.363223 | 97.959% +/- 0.039% | 689.31 | 380.3 |
| kvarn8-kvarn4 | 1656.00 | 40.4% | 0.002533 | 99.78% | 0.004515 | 0.007146 | 0.020414 | 0.086218 | 93.9% | 15.533092 | 97.867% +/- 0.040% | 645.67 | 406.0 |
| kvarn6-kvarn5 | 1528.00 | 37.3% | 0.002602 | 99.78% | 0.004247 | 0.006751 | 0.019168 | 0.079818 | 94.5% | 19.412773 | 98.021% +/- 0.038% | 692.77 | 378.4 |
| kvarn8-kvarn3 | 1528.00 | 37.3% | 0.003529 | 99.69% | 0.006129 | 0.009747 | 0.027324 | 0.121564 | 90.64% | 27.013851 | 97.623% +/- 0.042% | 649.84 | 403.4 |
| kvarn5-kvarn5 | 1400.00 | 34.2% | 0.002705 | 99.77% | 0.004366 | 0.006915 | 0.019942 | 0.083457 | 94.16% | 22.770369 | 97.912% +/- 0.039% | 699.8 | 374.6 |
| kvarn6-kvarn4 | 1400.00 | 34.2% | 0.002831 | 99.75% | 0.004553 | 0.007216 | 0.020422 | 0.091507 | 93.4% | 20.662098 | 97.865% +/- 0.040% | 694.79 | 377.3 |
| kvarn8-kvarn2 | 1400.00 | 34.2% | 0.009494 | 99.09% | 0.017839 | 0.027727 | 0.071455 | 0.325652 | 73.9% | 21.115437 | 96.120% +/- 0.053% | 651.45 | 402.4 |
| kvarn5-kvarn4 | 1272.00 | 31.1% | 0.002824 | 99.76% | 0.004707 | 0.007456 | 0.020926 | 0.093313 | 93.23% | 17.434282 | 97.856% +/- 0.040% | 700.73 | 374.1 |
| kvarn6-kvarn3 | 1272.00 | 31.1% | 0.003533 | 99.68% | 0.006181 | 0.009842 | 0.026777 | 0.123369 | 90.47% | 16.849087 | 97.678% +/- 0.042% | 697.01 | 376.1 |
| kvarn5-kvarn3 | 1144.00 | 27.9% | 0.003515 | 99.69% | 0.006348 | 0.009937 | 0.027889 | 0.118848 | 90.88% | 12.472089 | 97.547% +/- 0.043% | 701.67 | 373.6 |
| kvarn6-kvarn2 | 1144.00 | 27.9% | 0.009301 | 99.11% | 0.018093 | 0.028055 | 0.071621 | 0.310819 | 75.01% | 20.337669 | 96.115% +/- 0.053% | 697.56 | 375.8 |
| kvarn5-kvarn2 | 1016.00 | 24.8% | 0.009813 | 99.06% | 0.018174 | 0.028392 | 0.073407 | 0.344122 | 72.55% | 21.992970 | 96.156% +/- 0.053% | 705.26 | 371.7 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn8-kvarn8 | 2168.00 | 52.9% | 0.000550 | 99.95% | 0.000790 | 0.001213 | 0.003571 | 0.015574 | 98.86% | 8.929892 | 99.038% +/- 0.027% | 669.93 | 391.3 |
| kvarn8-kvarn6 | 1912.00 | 46.7% | 0.000525 | 99.96% | 0.000809 | 0.001240 | 0.003429 | 0.013725 | 99.05% | 10.572882 | 99.004% +/- 0.027% | 679.83 | 385.6 |
| kvarn8-kvarn5 | 1784.00 | 43.6% | 0.000570 | 99.95% | 0.000903 | 0.001389 | 0.003886 | 0.016143 | 98.81% | 8.564358 | 98.987% +/- 0.028% | 683.02 | 383.8 |
| kvarn6-kvarn6 | 1656.00 | 40.4% | 0.000543 | 99.96% | 0.000878 | 0.001341 | 0.003781 | 0.015912 | 98.83% | 8.530037 | 98.999% +/- 0.027% | 732.04 | 358.1 |
| kvarn8-kvarn4 | 1656.00 | 40.4% | 0.000729 | 99.94% | 0.001276 | 0.002009 | 0.005574 | 0.022666 | 98.17% | 8.352147 | 98.837% +/- 0.030% | 683.56 | 383.5 |
| kvarn6-kvarn5 | 1528.00 | 37.3% | 0.000697 | 99.94% | 0.000957 | 0.001490 | 0.004060 | 0.016759 | 98.75% | 14.932450 | 98.939% +/- 0.028% | 734.3 | 357.0 |
| kvarn8-kvarn3 | 1528.00 | 37.3% | 0.001579 | 99.85% | 0.002921 | 0.004734 | 0.012936 | 0.061181 | 94.46% | 7.800819 | 98.408% +/- 0.035% | 685.7 | 382.3 |
| kvarn5-kvarn5 | 1400.00 | 34.2% | 0.000692 | 99.94% | 0.001143 | 0.001791 | 0.005031 | 0.019015 | 98.52% | 10.226999 | 98.899% +/- 0.029% | 742.62 | 353.0 |
| kvarn6-kvarn4 | 1400.00 | 34.2% | 0.001017 | 99.91% | 0.001341 | 0.002104 | 0.005894 | 0.023520 | 98.08% | 14.792441 | 98.796% +/- 0.030% | 736.98 | 355.7 |
| kvarn8-kvarn2 | 1400.00 | 34.2% | 0.007060 | 99.31% | 0.014706 | 0.023042 | 0.059432 | 0.255041 | 77.81% | 7.684875 | 96.463% +/- 0.051% | 688.95 | 380.5 |
| kvarn5-kvarn4 | 1272.00 | 31.1% | 0.001047 | 99.91% | 0.001516 | 0.002420 | 0.006727 | 0.029455 | 97.5% | 10.981697 | 98.759% +/- 0.031% | 744.3 | 352.2 |
| kvarn6-kvarn3 | 1272.00 | 31.1% | 0.001619 | 99.85% | 0.002990 | 0.004821 | 0.013473 | 0.062179 | 94.36% | 7.625892 | 98.338% +/- 0.035% | 740.52 | 354.0 |
| kvarn6-kvarn2 | 1144.00 | 27.9% | 0.007148 | 99.3% | 0.014931 | 0.023303 | 0.060942 | 0.261539 | 77.31% | 9.853199 | 96.467% +/- 0.051% | 741.78 | 353.4 |
| kvarn5-kvarn3 | 1144.00 | 27.9% | 0.001754 | 99.83% | 0.003155 | 0.005106 | 0.014005 | 0.066446 | 93.96% | 8.938780 | 98.345% +/- 0.035% | 746.42 | 351.2 |
| kvarn5-kvarn2 | 1016.00 | 24.8% | 0.007140 | 99.30% | 0.014991 | 0.023365 | 0.060848 | 0.254938 | 77.82% | 6.150640 | 96.476% +/- 0.051% | 748.34 | 350.3 |
| Cache | KV cache (MiB) | Size vs bf16 | Mean KLD | Precision vs bf16 | 90% KLD | 95% KLD | 99% KLD | 99.9% KLD | 99.9% precision vs bf16 | Maximum KLD | Same top p | Tok/s | Elapsed (s) |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| kvarn8-kvarn8 | 4312.00 | 52.6% | 0.000387 | 99.96% | 0.000780 | 0.001201 | 0.003289 | 0.012994 | 98.71% | 0.446070 | 99.064% +/- 0.027% | 501.14 | 523.1 |
| kvarn8-kvarn6 | 3800.00 | 46.4% | 0.000395 | 99.96% | 0.000802 | 0.001221 | 0.003287 | 0.012882 | 98.73% | 0.368090 | 99.038% +/- 0.027% | 509.31 | 514.7 |
| kvarn8-kvarn5 | 3544.00 | 43.3% | 0.000437 | 99.96% | 0.000871 | 0.001334 | 0.003561 | 0.014522 | 98.56% | 0.463152 | 99.024% +/- 0.027% | 510.6 | 513.4 |
| kvarn6-kvarn6 | 3288.00 | 40.1% | 0.000423 | 99.96% | 0.000861 | 0.001321 | 0.003609 | 0.015358 | 98.48% | 0.261352 | 99.051% +/- 0.027% | 566.55 | 462.7 |
| kvarn8-kvarn4 | 3288.00 | 40.1% | 0.000575 | 99.94% | 0.001210 | 0.001876 | 0.004963 | 0.019958 | 98.03% | 0.479687 | 98.863% +/- 0.029% | 512.2 | 511.8 |
| kvarn6-kvarn5 | 3032.00 | 37% | 0.000448 | 99.96% | 0.000916 | 0.001414 | 0.003815 | 0.014024 | 98.61% | 0.508397 | 98.965% +/- 0.028% | 570.75 | 459.3 |
| kvarn8-kvarn3 | 3032.00 | 37% | 0.001211 | 99.88% | 0.002616 | 0.004163 | 0.010882 | 0.045876 | 95.52% | 0.551635 | 98.496% +/- 0.034% | 515.12 | 508.9 |
| kvarn5-kvarn5 | 2776.00 | 33.9% | 0.000537 | 99.95% | 0.001094 | 0.001694 | 0.004521 | 0.018230 | 98.2% | 0.837041 | 98.907% +/- 0.029% | 576.65 | 454.6 |
| kvarn6-kvarn4 | 2776.00 | 33.9% | 0.000597 | 99.94% | 0.001236 | 0.001911 | 0.004967 | 0.019079 | 98.12% | 1.141340 | 98.863% +/- 0.029% | 571.12 | 459.0 |
| kvarn8-kvarn2 | 2776.00 | 33.9% | 0.005841 | 99.42% | 0.013071 | 0.020548 | 0.050159 | 0.191198 | 82.6% | 1.809292 | 96.765% +/- 0.049% | 515.93 | 508.1 |
| kvarn5-kvarn4 | 2520.00 | 30.8% | 0.000670 | 99.93% | 0.001407 | 0.002189 | 0.005673 | 0.024618 | 97.57% | 0.606157 | 98.816% +/- 0.030% | 578.3 | 453.3 |
| kvarn6-kvarn3 | 2520.00 | 30.8% | 0.001209 | 99.88% | 0.002633 | 0.004197 | 0.010818 | 0.044875 | 95.62% | 0.984286 | 98.503% +/- 0.034% | 573.24 | 457.3 |
| kvarn5-kvarn3 | 2264.00 | 27.6% | 0.001293 | 99.87% | 0.002798 | 0.004503 | 0.011862 | 0.044334 | 95.67% | 2.241802 | 98.427% +/- 0.034% | 581.25 | 451.0 |
| kvarn6-kvarn2 | 2264.00 | 27.6% | 0.005834 | 99.42% | 0.013002 | 0.020420 | 0.051026 | 0.191961 | 82.54% | 1.639544 | 96.806% +/- 0.049% | 575.51 | 455.5 |
| kvarn5-kvarn2 | 2008.00 | 24.5% | 0.005910 | 99.41% | 0.013283 | 0.020801 | 0.051606 | 0.185401 | 83.08% | 1.768032 | 96.737% +/- 0.049% | 583.71 | 449.1 |
The first KVarN run already had the important low-bit result: kvarn4-kvarn4 was not merely a better q4 row, it reached the normal q5 tier while using less memory than ordinary q4_0. The new rows repeat that shape higher up the ladder. The symmetric rows prove the shift cleanly, and the asymmetric pairs are the ones with highest value per MiB.
There are two different questions: one is whether KVarN can match a higher ordinary quant tier at lower memory, and the other is which pairs you should actually run. kvarn5-kvarn5 is the clean q6-class proof, but kvarn5-kvarn4 is the sharper mid-tier value row. kvarn6-kvarn6 is the conservative q8-class proof, while kvarn6-kvarn5 is a more aggressive variant.
The clean read is that KVarN shifts the useful cache ladder down by about one ordinary memory tier. This is not true for every possible K/V split. It holds for balanced rows and for the one-bit asymmetric rows that keep K slightly higher than V. The table below separates proof rows from value rows so the result does not get flattened into a symmetric-only story.
| Role | KVarN row | KVarN size | Usual quant tier reached | Usual size | Concrete evidence |
|---|---|---|---|---|---|
| Clean q5-class proof | kvarn4-kvarn4 | 27.9% | q5_0 | 34.4% | On Q5_K_S 64k, mean/tail KLD is 0.002974 / 0.094819 vs q5_0 at 0.003206 / 0.099073. |
| Aggressive q4/q5 value | kvarn4-kvarn3 | 24.8% | Above q4_0, approaching q5_0-q4_0 | 28.1% / 31.3% | On Q5_K_S 64k, it beats q4_0 on mean KLD, 0.003824 vs 0.004711, while using less memory than both q4_0 and q5_0-q4_0. The tail is weaker than q4_0, so this is a value row, not a clean q5 proof. |
| Clean q6-class proof | kvarn5-kvarn5 | 34.2% | q6_0 | 40.6% | On both IQ4_XS rows it beats q6_0 on mean and 99.9% KLD. On Q5_K_S 64k it loses mean narrowly, 0.002705 vs 0.002614, but wins the tail, 0.083457 vs 0.090800. |
| Mid-tier value row | kvarn5-kvarn4 | 31.1% | q5_1 / q5_0 | 37.5% / 34.4% | On Q5_K_S 64k, it beats q5_1 on both mean and tail, 0.002824 / 0.093313 vs 0.002911 / 0.098354, while using about 6.4 percentage points less cache. |
| Clean q8-class proof | kvarn6-kvarn6 | 40.4% | q8_0 / q8_0-q6_0 | 53.1% / 46.9% | On Q5_K_S 64k, it is effectively tied with symmetric q8_0 and strictly beats q8_0-q6_0. On IQ4_XS 64k, it beats both. |
| Aggressive q8-ish value | kvarn6-kvarn5 | 37.3% | q8-ish / optimized q8-q5 tier | 45.3% to 53.1% | On IQ4_XS 128k, it beats symmetric q8_0, 0.000448 / 0.014024 vs 0.000482 / 0.014951, at 37.0% instead of 53.1%. On Q5_K_S 64k it is a tradeoff against q8_0-q6_0: worse mean, better tail, but much smaller cache. |
| Ceiling check | kvarn8-kvarn8 | 52.9% | q8_0 | 53.1% | It lands in the same q8 band, but the memory is also q8-sized, so the value row has already happened at kvarn6. |
KVarN is not just better within its own nominal bit width. In the useful range, it often gives the quality tier above the memory tier it occupies. The row a user picks, however, may be one step asymmetric rather than symmetric.
kvarn5-kvarn5 should not be sold as q8-class, but it is the first row that looks q6-class while staying in the q5 memory band. The Q5_K_S 64k row is mixed because mean KLD is slightly worse than q6_0, 0.002705 vs 0.002614, but the 99.9% tail is better, 0.083457 vs 0.090800.
| Bench | q6_0 | kvarn5-kvarn5 | Read |
|---|---|---|---|
Q5_K_S 64k | 40.6%, 0.002614 / 0.090800 | 34.2%, 0.002705 / 0.083457 | Mean loses narrowly, tail wins. |
IQ4_XS 64k | 40.6%, 0.000766 / 0.020407 | 34.2%, 0.000692 / 0.019015 | Strictly beats q6 and is cheaper. |
IQ4_XS 128k | 40.6%, 0.000589 / 0.019175 | 33.9%, 0.000537 / 0.018230 | Strictly beats q6 and is cheaper. |
The more practical mid-tier row is kvarn5-kvarn4. On the main Q5_K_S row, it uses 31.1% of the bf16 KV cache and scores 0.002824 / 0.093313. That is better than q5_1 at 37.5%, better than q5_0 at 34.4%, and better than the old asymmetric q5_0-q4_0 at 31.3%. Overall, it's a very solid mid-tier recommendation.
| Bench | kvarn5-kvarn5 | kvarn5-kvarn4 | Read |
|---|---|---|---|
Q5_K_S 64k | 34.2%, 0.002705 / 0.083457 | 31.1%, 0.002824 / 0.093313 | 5/5 is cleaner; 5/4 is the stronger value row. |
IQ4_XS 64k | 34.2%, 0.000692 / 0.019015 | 31.1%, 0.001047 / 0.029455 | 5/4 still beats ordinary q5_0 at lower memory. |
IQ4_XS 128k | 33.9%, 0.000537 / 0.018230 | 30.8%, 0.000670 / 0.024618 | Same pattern: 5/5 for q6-class, 5/4 for the cheaper middle. |
So kvarn5 has two jobs. kvarn5-kvarn5 proves the q6-at-q5-memory claim. kvarn5-kvarn4 is the row I would highlight when talking about value per size, because it sits around the old q5_0-q4_0 memory band while beating ordinary q5.
The q8 comparison starts to become real at kvarn6. Against symmetric q8_0, kvarn6-kvarn6 is effectively tied on the main Q5_K_S row: 0.002338 / 0.078797 vs 0.002328 / 0.078709, but it comes at 40.4% of bf16 KV instead of 53.1%.
| Bench | Conservative KVarN6 row | Aggressive KVarN6 row | Read |
|---|---|---|---|
Q5_K_S 64k | kvarn6-kvarn6: 40.4%, 0.002338 / 0.078797 | kvarn6-kvarn5: 37.3%, 0.002602 / 0.079818 | 6/6 strictly beats q8_0-q6_0; 6/5 trades mean for tail and size. |
IQ4_XS 64k | kvarn6-kvarn6: 40.4%, 0.000543 / 0.015912 | kvarn6-kvarn5: 37.3%, 0.000697 / 0.016759 | 6/6 is the clean q8-class row; 6/5 still beats q8_0-q5_1 cheaper. |
IQ4_XS 128k | kvarn6-kvarn6: 40.1%, 0.000423 / 0.015358 | kvarn6-kvarn5: 37.0%, 0.000448 / 0.014024 | 6/5 is the sharper row here, beating symmetric q8_0 at much lower memory. |
Against optimized usual quant rows, kvarn6-kvarn6 is the clean claim. It beats q8_0-q6_0 on Q5_K_S 64k and IQ4_XS 64k while using about 40% instead of 47%. kvarn6-kvarn5 is not as clean on the main Q5 row because its mean KLD is worse than q8_0-q6_0, 0.002602 vs 0.002499, but its 99.9% tail is better, 0.079818 vs 0.081616, and it uses 37.3% instead of 46.9%. That makes it a great aggressive value row on the high end.
kvarn8-kvarn8 is sort of meaningless. On Q5_K_S 64k, it scores 0.002361 / 0.076809, while ordinary q8_0 scores 0.002328 / 0.078709. On IQ4_XS 128k, it scores 0.000387 / 0.012994 against q8_0 at 0.000482 / 0.014951. Those are q8-class numbers, and sometimes better than q8, but the footprint is also q8-class: 52.9% at 64k and 52.6% at 128k.
The asymmetric kvarn8 rows do not change the recommendation much. kvarn8-kvarn5 can beat ordinary q8_0 on the IQ4 rows at lower memory, but kvarn6-kvarn6 and kvarn6-kvarn5 already reach the same useful tier with less cache. Once kvarn6 has reached q8-class fidelity, extra K bits mostly become a diagnostic check rather than a preset.
That makes kvarn8 useful mostly as a ceiling check. It shows that the KVarN path can sit at the top of the measured distribution-fidelity range, but it also shows why that top row is not necessary. If kvarn6-kvarn6 already reaches the q8 tier at roughly q6 memory, kvarn8-kvarn8 mostly answers a diagnostic question rather than a preset question.
The follow-up also puts a limit on the old K-first rule. K still deserves priority when there is one extra bit to spend, but dumping bits into K while starving V stops working quickly. At the same 34.2% footprint on Q5_K_S 64k, kvarn5-kvarn5 scores 0.002705 / 0.083457, kvarn6-kvarn4 slips to 0.002831 / 0.091507, and kvarn8-kvarn2 collapses to 0.009494 / 0.325652. More K does not rescue a bad V side.
| Same-size comparison | Better balanced row | More K-heavy row | Read |
|---|---|---|---|
40.4% on Q5_K_S | kvarn6-kvarn6: 0.002338 / 0.078797 | kvarn8-kvarn4: 0.002533 / 0.086218 | Balanced wins. |
34.2% on Q5_K_S | kvarn5-kvarn5: 0.002705 / 0.083457 | kvarn6-kvarn4: 0.002831 / 0.091507 | One extra K bit is not free. |
31.1% on Q5_K_S | kvarn5-kvarn4: 0.002824 / 0.093313 | kvarn6-kvarn3: 0.003533 / 0.123369 | A two-bit K/V gap is already too much. |
27.9% on Q5_K_S | kvarn4-kvarn4: 0.002974 / 0.094819 | kvarn5-kvarn3: 0.003515 / 0.118848 | The old symmetric row still beats the higher-K split. |
The practical rule is simple: K can take the extra bit, but V should stay within one bit unless the row has data proving otherwise. The useful asymmetric rows are near-balanced, especially kvarn4-kvarn3, kvarn5-kvarn4, and kvarn6-kvarn5. The rows with a two-bit or larger gap are mostly negative evidence.
With the follow-up included, the preset ladder changes more than the original 2/3/4-bit article suggested. The old ladder had kvarn4-kvarn4 at the top because that was the largest tested KVarN row. The higher-bit sweep makes kvarn4-kvarn4 a memory-saving row, not the quality ceiling.
| Preset | Size vs bf16 | Quality class | What it is for |
|---|---|---|---|
bf16 / bf16 | 100.0% | Reference | Blame-isolation and full-quality baseline. |
kvarn8-kvarn8 | 52.9% | q8 ceiling | Diagnostic row; not the main value preset. |
kvarn6-kvarn6 | 40.4% | clean q8-class | Conservative high-end KVarN preset. |
kvarn6-kvarn5 | 37.3% | aggressive q8-ish | Use when kvarn6-kvarn6 misses the fit and a small mean-KLD tradeoff is acceptable. |
kvarn5-kvarn5 | 34.2% | q6-class | Clean q6-at-q5-memory proof row. |
kvarn5-kvarn4 | 31.1% | q5/q6 middle | Probably the best mid-tier value preset. |
kvarn4-kvarn4 | 27.9% | q5-class | q5-ish quality below ordinary q4 memory. |
kvarn4-kvarn3 | 24.8% | above q4_0 | Below-q4 memory without falling to turbo3 quality. |
kvarn3-kvarn3 | 21.7% | compact | Better turbo3-class option if memory is tight. |
kvarn2-kvarn2 | 15.4% | emergency compression | Useful endpoint, not a normal recommendation. |
In these runs, KVarN shifts the practical ladder down by roughly one memory tier: q5 quality at q4 memory, q6 quality at q5 memory, and q8-class quality around q6 memory. The symmetric rows show the ladder, and the one-bit asymmetric rows are often where the best value per size sits. kvarn8 reaches the ceiling, but kvarn6 is the reason the ceiling matters.
The graphs below plot mean and 99.9% KL divergence against KV-cache size, with separate plots for the full range and the sub-4-bit range. Each point is one cache configuration from the tables above.



