Unveiling the Potential of Quantization with MXFP4: Strategies for Quantization Error Reduction
Abstract
Large Language Models (LLMs) have intensified the need for low-precision formats for efficient inference. The Open Compute Project Microscaling (MX) standard is attractive due to its favorable hardware efficiency, but its 4-bit variant (MXFP4) lags behind NVIDIA’s NVFP4 in accuracy, limiting adoption. We introduce two software-only techniques, Overflow-Aware Scaling (OAS) and Macro Block Scaling (MBS), that improve MXFP4 quantization fidelity without requiring hardware changes. OAS reduces overall errors by increasing effective dynamic range under power-of-two block scaling, while MBS allocates higher-precision scaling at a coarser granularity to better preserve outliers. Across multiple LLMs and standard downstream benchmarks, OAS and MBS reduce the end-to-end accuracy gap between MXFP4 and NVFP4 from about 10% to below 1% on average, while incurring modest GEMM overhead (6.2% on average). These results re-establish MXFP4 as a practical alternative to NVFP4, enabling near-NVFP4 accuracy while retaining MX’s hardware-efficiency advantages (e.g., 12% relative area savings in tensor cores).
Lay Summary
Modern Large Language Models are expensive to run, both in cost and energy. One way to reduce that cost is to compress their internal numbers, storing each with fewer bits than the original precision. The industry has converged on two competing 4-bit number formats. NVIDIA's NVFP4 is more accurate while the alternative MXFP4 is cheaper in terms of hardware cost but loses noticeable accuracy in practice. We pinpointed two specific reasons MXFP4 was lagging, and designed two software-only techniques. The first, Overflow-Aware Scaling, chooses each value's scaling factor more carefully so fewer extreme numbers get clipped. The second, Macro Block Scaling, adds a tiny higher-precision correction shared across larger groups of elements to better capture the rare large values that matter most. Together, these techniques close almost all of MXFP4's accuracy gap from about 10% down to 1%. These results re-establish MXFP4 as a practical alternative to NVFP4, enabling near-NVFP4 accuracy while retaining MX’s hardware-efficiency advantages.