FPTQuant: Function-Preserving Transforms for LLM Quantization
Abstract
Large language models (LLMs) require substantial compute, and thus energy, at inference time. While quantizing weights and activations is effective at improving efficiency, naive quantization of LLMs can significantly degrade performance due to large magnitude outliers. This paper describes FPTQuant, which introduces three novel, lightweight, and expressive function-preserving transforms (FPTs) to facilitate quantization of transformers: (1) a mergeable pre-RoPE transform for queries and keys, (2) a mergeable transform for values, (3) a cheap, dynamic scaling transform. By leveraging the equivariances and independencies inherent to canonical transformer operation, we designed these FPTs to maintain the model’s function while shaping the intermediate activation distributions to be more quantization friendly. FPTQuant requires no custom kernels and adds virtually no overhead during inference. The FPTs are trained both locally to reduce outliers, and end-to-end such that the outputs of the quantized and full-precision models match. FPTQuant enables static INT4 quantization with minimal overhead and shows SOTA speed-up of up to 3.9x over FP. Empirically, FPTQuant has an excellent accuracy-speed trade-off—it is performing on par or exceeding most prior work and only shows slightly lower accuracy compared to a method that is up to 29% slower.
Lay Summary
Modern AI language models need a lot of computing power—and energy—when they generate responses. One common way to make them faster and more efficient is called quantization, which simplifies the numbers the model uses. However, doing this in a straightforward way can harm performance because some values inside the model are unusually large and do not compress well. This paper introduces a new method called FPTQuant that makes quantization work much better for these models. It does this by adding a few small adjustments (called function-preserving transforms) inside the model. These adjustments reshape the internal data so that the model becomes easier to compress—without changing how it behaves or what outputs it produces. The resulting model gives outputs that are close to the original model, but is up to 3.9 times faster.