Approximation Error Upper and Lower Bounds for Hölder Class with Transformers
Abstract
Lay Summary
Why is the Transformer — the core architecture of modern AI models — so powerful? And what scale must a Transformer reach to control errors within a specific range when processing complex tasks? To answer these questions, we analytically evaluated the expressivity of Transformers through specific function approximation tasks, and aimed to establish a rigorous mathematical explanation for their success. In this paper, we derived the minimum and maximum model scale required for a standard yet mathematically simplified Transformer to achieve a given target accuracy. From another perspective, for any given model size, our results reveal the absolute best- and worst-case theoretical accuracy limits that a Transformer can reach. These theoretical guarantees provide solid mathematical backing for Transformers' remarkable performance. Furthermore, by extending our findings to general regression tasks, we offer a theoretical explanation for why Transformers perform so well in real-world applications.