Probing the Inductive Bias of Neural Networks through Learning Random Cellular Automata
Abstract
Lay Summary
Neural networks learn structures in real-world data remarkably well, spanning all kinds of natural patterns. This can only be possible if neural networks are already pre-tuned to a strong statistical pattern that is common to all of these phenomena. In machine learning, this is called prior-information. At this point, we have limited insight into what prior information networks use exactly; we only know that the structure of how they are built encodes this information. A common hypothesis is that the fundamental rules of physics might already dictate this prior information. But despite our understanding of physics, we do not know which aspects might be essential. This paper follows a simple idea: It tries various variants of artificial physics with certain aspects (such as locality of interactions, chaotic dynamics, etc) turned on and off and probes whether typical neural network architectures can learn the patterns obtained. The result is that common physical principles are helpful, some even essential, but not sufficient to explain learnability. An additional ingredient we find is the complexity of the pattern: Patterns where many parts interact closely become harder to learn the more parts need to be considered simultaneously. Measuring this predicts with surprising accuracy whether the neural network can learn the pattern. While such a complexity limit is not directly a law of physics, it might possibly arise as emergent property of long-running dynamical processes, which at this point still remains hypothetical.