Towards Atoms of Large Language Models
Abstract
Lay Summary
Large language models are powerful, but we still do not know what their basic “pieces of meaning” are. This makes it hard to understand how they store ideas, facts, or concepts inside the model. We propose Atom Theory, a way to find these basic pieces, which we call atoms. A good atom should do two things: it should accurately explain the model’s internal activity, and it should stay clear when the model sees different text. We show that commonly studied units, such as neurons and learned features, do not fully satisfy both requirements. We then develop a method to find better atoms in Gemma and Llama models. These atoms are much more accurate, stable, and easier to connect to single human-understandable ideas. This work gives a clearer way to look inside large language models and may help make them easier to understand, analyze, and trust.