REAL-Q brings gradient descent to LLM quantization
A new post-training quantization method, REAL-Q, uses block-wise gradient descent to reduce end-to-end KL divergence by up to about 49% compared with second-order methods on LLaMA-3.1 and Qwen3 at W4A16. The result matters for resource-constrained de...