Tailoring the Quantization Space
for 1-Bit KV Cache Compression

1 Seoul National University · 2 Stanford University

* Equal contribution   † Corresponding author

TaSQ (Tailored Space for Vector Quantization) is an LLM KV cache vector quantization (VQ) method that tailors the quantization space for accurate ultra-low-bit compression. Through query-guided channel weighting and covariance-aware grouping, TaSQ makes the VQ target space better reflect channel-wise sensitivity and inter-channel dependencies. Consequently, TaSQ substantially reduces quantization error in the 1-bit regime, outperforming existing baselines across general, reasoning, and long-context retrieval benchmarks. Our SGLang implementation supports up to 14× larger batch sizes and achieves 1.87× higher peak throughput compared to the BF16 baseline.

Motivations

Four motivation panels showing channel-wise query activations, inter-channel key correlations, pre- and post-RoPE key distributions, and reconstruction error across layers.

Figure 1. Key-channel structure in Llama-3.1-8B: (a) query activation ranges, (b) absolute key correlations, (c) key distributions before and after RoPE, and (d) VQ reconstruction error across layers. All panels use the same 2,048-token GPQA window; (a–c) show layer 15, KV head 0.

1. Sensitivity of key channels to queries. Quantization errors in different key channels affect attention scores differently depending on the corresponding query activations. Figure 1a shows that query distributions vary substantially across channels, implying unequal sensitivity of key channels to reconstruction error. This is misaligned with the objective of Euclidean VQ, which treats all channels equally.

2. Inter-channel correlation. Key channels exhibit non-uniform dependencies, as shown in Figure 1b: some channel pairs are strongly correlated, while others are nearly independent. Since VQ represents multiple channels jointly with a shared codebook, its ability to exploit these dependencies depends on which channels are grouped together.

3. Pre-RoPE vs. post-RoPE keys. RoPE applies position-dependent rotations that spread the relatively compact pre-RoPE key distribution, as shown in Figure 1c. This makes post-RoPE keys harder to represent with a shared codebook. Consistently, Figure 1d shows that pre-RoPE VQ achieves 35% lower total reconstruction error.

Method

TaSQ transforms pre-RoPE keys using query-guided weighting, shared-scale normalization, and covariance-aware channel grouping before vector quantization.

Figure 2. Overview of TaSQ.

The figure above illustrates the overall pipeline of TaSQ. Three transformations are applied sequentially to tailor the pre-RoPE key space, and the effect of each step is visualized in the figure.

1. Query-guided Channel Weighting. We derive channel weights from query–key dot-product error and rescale key channels accordingly. Euclidean reconstruction error in the weighted space then serves as a surrogate for expected squared attention-logit error.

2. Cross-head Shared-scale Normalization. We normalize each token using a single scale shared across all KV heads. This suppresses token-level magnitude outliers while keeping normalization metadata small.

3. Covariance-aware Channel Grouping. We group dependent channels into the same VQ codebook using a covariance-based criterion. Groups preserve complete RoPE pairs, allowing RoPE to remain efficient at inference time.

Weighting and grouping are absorbed into projection weights and decoder codebooks offline, while codeword lookup, scale restoration, RoPE, and the query–key dot product are fused into the attention kernel. Values use standard VQ with contiguous channel groups.

Experimental Results

Benchmark
Model Method Bits (K/V) GSM8K MATH500 MBPP HumanEval BBH MMLU Avg.
Llama-3.1-8B-Instruct BF16 16.000/16.000 83.62 42.20 59.60 62.20 73.38 62.64 63.94
Llama-3.1-8B-Instruct CQ 1.250/1.250 68.01 19.60 52.00 54.27 40.08 54.21 48.03
Llama-3.1-8B-Instruct NovaKV 1.375/1.250 76.95 29.40 55.00 53.66 48.46 56.43 53.32
Llama-3.1-8B-Instruct NSNQuant 1.238/1.238 73.39 31.40 50.20 57.93 48.57 58.33 53.30
Llama-3.1-8B-Instruct TaSQ 1.266/1.250 81.73 34.00 58.40 58.54 64.28 58.33 59.21
Qwen3-4B BF16 16.000/16.000 86.28 72.60 64.00 81.71 78.03 74.72 76.22
Qwen3-4B CQ 1.250/1.250 70.36 60.80 45.40 65.85 52.08 64.14 59.77
Qwen3-4B NovaKV 1.375/1.250 84.53 66.80 62.60 76.22 68.80 67.21 71.03
Qwen3-4B NSNQuant 1.238/1.238 65.58 52.40 46.20 67.68 49.57 63.29 57.45
Qwen3-4B TaSQ 1.266/1.250 85.44 69.80 62.60 77.44 68.28 71.26 72.47


TaSQ achieves the best overall performance among the evaluated quantized methods across general, reasoning, and long-context benchmarks. Notably, on a long-context benchmark, its advantage becomes increasingly pronounced as the context length grows.

Serving Efficiency

Throughput by batch size, prefill latency, and generated-token distributions comparing TaSQ with baselines.

Figure 3. Qwen3-4B-Thinking-2507 on one RTX 6000 Ada GPU. Throughput uses 2,048-token prompts and up to 32,768 generated tokens. Prefill latency is measured at batch size 1, and generation lengths are measured on LiveCodeBench-v6.

Throughput / TTFT. TaSQ introduces little additional serving overhead over CQ, a minimal KV cache VQ baseline without additional runtime transformations, as shown in Figure 3a. Compared with BF16, TaSQ supports up to 14× larger batches and achieves 1.87× higher peak throughput. Although VQ encoding increases time-to-first-token by 10–14% for 8k–32k prompts, this one-time prefill cost is amortized over long generations.

Reasoning stability. Aggressive KV cache compression can destabilize reasoning models, leading to repetitive or unterminated generations. In Figure 3c, TaSQ keeps generation lengths and cap-hit rates close to BF16, while CQ more frequently reaches the 32k generation limit, wasting decoding budget on failed outputs.