A geometric relation of the error introduced by sampling a language model's output distribution to its internal state
Abstract
Lay Summary
AI chatbots work by generating one word at a time. They do this by assigning probabilities to each possible word at every time step. At most time steps, the chatbots are confident, and so only one word out of the many possible ones is very probable. However, there are some time steps where this is not the case, and the chatbot is conflicted between a few possible next words. These conflicted moments can be very significant, changing the overall answer to a reasoning problem, despite corresponding to only a small change between two similar words. The confidence of the chatbot is usually summarised in a single number, but we are able to show that there is more structure here. By making a particular geometric construction, we are able to relate these uncertain moments to how the chatbot internally conceptualises high-level strategic elements of its task. We use ordinary commercial AI chatbots and make them think through chess positions to analyse this.