系列:Operator Explainers

Operator Explainers (2): Model Quantization for Deployment — the precision-compression spectrum from FP16 to FP4

0. The One-Line Thread

Model quantization is the art of trading precision for memory/bandwidth/compute — it is orthogonal to distillation (which trades capacity for size); the two stack (distill a small model, then quantize it for deployment). This post consolidates the FP8/INT4/FP4/KV-Cache quantization notes that were scattered across daily tracking into one spectrum, landing on real vLLM/SGLang and speculative-decoding deployments.

1. Precision Spectrum and Compression

PrecisionBitsvs FP16Typical lossArena
FP16 / BF16161×baselinetraining, high-fidelity inference
FP8 (E4M3 / E5M2)82×~0.3%weights+acts, KV Cache
INT882×slightly above FP8edge, older HW
INT4 (mixed)44×~0.5%on-device / onboard
FP444×needs QATnext-gen train/infer

2. Post-Training Quantization (PTQ): compression without gradients

PTQ quantizes a trained model directly; the core question is “how to assign scaling factors”:

Rotation-based quantization (outlier suppression) — the frontier branch of PTQ. Its mathematical lever is the rotational invariance of linear maps: for any orthogonal R (R⁻¹=Rᵀ), XW = (XR⁻¹)(RW). Left-multiplying weights by R and right-multiplying activations by Rᵀ leaves the product unchanged but “stirs” the numerical distribution — the activation outliers concentrated in a few channels get flattened, so 4-bit quantization no longer overflows its dynamic range. That is the key that makes W4A4 viable.

  • QuaRot (NeurIPS 2024): uses a randomized Hadamard transform to build the orthogonal rotation, applying R1–R4 rotations across weights, activations, and KV Cache simultaneously — the first to hit W4A4KV4 (4-bit weights + 4-bit acts + 4-bit KV) near-losslessly on LLaMA-2/3, training-free and plug-and-play.
  • SpinQuant (arXiv:2405.16406): upgrades the rotation from “random fixed” to “learnable” — optimizing the rotation matrices on the Stiefel manifold via Cayley SGD so the rotation fits both weight and activation distributions. On LLaMA-3 it beats QuaRot across the board, recovering up to 45.1% more accuracy, and folds RoPE into the rotation optimization too.
  • DuQuant: adds channel permutation on top of rotation, scattering outliers into adjacent channels for even steadier W4A4.

Deployment note: rotation is orthogonal to AWQ/GPTQ — AWQ protects salient channels, GPTQ does second-order compensation, the rotation camp flattens outliers; the three stack (rotate then AWQ), and that stacking is the key piece for pushing 4-bit on-device inference to usable today.

VLA in practice: OpenVLA INT4 ≈ 4GB with near-lossless accuracy runs on Jetson Orin, leaving headroom for perception/planning/control — the only way “big models enter the physical world.”

3. Quantization-Aware Training (QAT): put the error into backprop

PTQ loss grows at INT4/FP4, so QAT is needed:

4. KV Cache quantization and eviction: the long-context memory lifeline

The main bottleneck of long-context inference is not FLOPs but KV Cache memory. Two orthogonal paths:

5. Landing on inference engines and speculative decoding

6. Deployment checklist

  1. Locate the bottleneck first: memory wall → KV quant/evict; bandwidth wall → INT4/AWQ weights; compute wall → FP8 forward.
  2. Prefer FP8 (2×, ~0.3%) ; drop to INT4 (AWQ, 4×) only on-device.
  3. FP4 must come with QAT, or only for drop-insensitive scenarios.
  4. Long text → E5M2; code/math → E4M3.
  5. Below INT2 is basically unusable — don’t trade usability for a headline number.

Tying it together: distillation shrinks capacity, quantization compresses precision, KV eviction cuts length, tiered caching (HiSparse) puts all three into one memory budget — that is the full toolbox for squeezing 70B+ models into single-card long-context serving today.

Note: this post consolidates the 2026-07-18 (KV Cache quant+evict), 2026-08-01 (FP8/INT4 deploy), 2026-08-04 (FP4 QAT) daily technical tracks into a systematic version.

觉得有用?欢迎点赞、收藏,或请作者喝咖啡 ☕️

支付宝收款码

支付宝

微信收款码

微信

💬 留言

评论由 Giscus 驱动(基于 GitHub Discussions)。 当前仓库 NaphJohn/LLM-blog 尚未启用 Discussions:请在 GitHub 仓库 Settings → General → Features 勾选 Discussions 后刷新本页,评论区即自动显示。