Fault Robustness of Custom Floating-Point and Integer Formats: Datatype Selection as a Reliability-Aware Compression Decision
Ramamoorthi Sreenivasan Haripriya ⋅ Jaynarayan T Tudu
Abstract
Deploying neural networks on resource-constrained edge devices requires both memory-efficient quantization and robustness to hardware-induced faults. We present a unified fault robustness study across fourteen low-precision datatypes spanning custom Float16, Float8, and integer representations under random multi-bit DRAM disturbance faults. Robustness is evaluated through top-1 accuracy degradation on MobileNetV2, EfficientNet-B0, and ResNet-18 across CIFAR-10, CIFAR-100, and Tiny ImageNet. Our results show that floating-point fault sensitivity is dominated by exponent-field width, with wider exponents causing severe error amplification under bit flips. Among all evaluated formats, E4M11 consistently provides the best trade-off between clean accuracy and fault robustness: it incurs only a $0.23\%$ average clean-accuracy drop relative to FP32 — on par with FP16 and BF16 — yet limits worst-case exponent-fault accuracy loss to $6.1\%$ on average, versus $34.6\%$ for FP16 and $36.2\%$ for BF16, a $5.7{\times}$/$6.0{\times}$ robustness gain. Although INT16 is structurally immune to exponent faults, it delivers ${\sim}10.7$\,dB lower Signal-to-Quantisation-Noise Ratio (SQNR) than E4M11 with no compression benefit; INT8 avoids exponent amplification but suffers ${\sim}58.8$\,dB lower SQNR and larger accuracy degradation on harder tasks. E4M11 thus uniquely combines near-FP32 accuracy, the highest SQNR (${\sim}79.5$\,dB) among all formats studied, and superior fault tolerance — establishing datatype selection as a zero-cost reliability lever requiring no retraining or error-correction hardware.
Chat is not available.
Successful Page Load