Modelling Attention with Aitchison Geometry: Token Distinguishability and Temperature Scaling
Abstract
Lay Summary
Long data passages are a challenge for large language models or AI models because they must decide which parts of the input are most important for the output. These "important scores" become much smaller as the data gets longer, making it difficult to tell whether the model is making useful choices or losing focus on the important pieces of information. We study this problem by measuring the importance scores relative to each other. This helps tell the difference between the scores shrinking because of the data's size or because the model is no longer making clear choices. We show that, in idealistic conditions, these relative differences do not disappear just because the input is long. However, in non-idealistic, real-world conditions, we show the strength of the adjustment to the scores needed to help the model make clear choices. We also show that measuring the relative scores provides important information in the real-world development and deployment of AI models.