Low-dimensional topology of deep neural networks
Abstract
Lay Summary
Life ties itself in knots: proteins fold into tangled shapes, and vaccine mRNA coils into knot-like loops. Astonishingly, today's artificial intelligence faces the very same problem. The networks behind chatbots and image recognizers must tell one kind of data from another. Yet engineers found their best tricks long before they understood why. Picture two classes of data as interlocking rings. We prove a plain network can never pull them apart, no matter how deep it goes or how long it trains. The fixes are modern AI's own hallmarks: attention inside ChatGPT, the shortcut "residual connections" in deep networks, and newer "folding" activations. Together they flip space like a loop of string, until the rings come free. Our experiments confirm it, and we even find these tangles hiding in real photos. The lesson is a design principle: winning architectures fold and untangle the shapes hidden in data.