Interpretability Should Prioritise Use-Inspired Basic Research for AI Safety
Abstract
There is currently a debate in the field of AI interpretability about whether pragmatic, incremental research which applies simple techniques is a more promising strategy for AI safety applications than ambitious, riskier research which builds new tools, theory and analysis to explain neural networks. This debate closely mirrors the debate in Innovation Economics about how to balance basic science, aiming at understanding, versus applied science, aiming at use, for increasing the rate of technological innovation. We follow Donald Stokes’ approach to this conflict in proposing the synthesis of Use-Inspired Basic Research, an approach in which researchers focus on practical problems, yet aim to deeply understand the relevant phenomena. Ideally Use-Inspired Basic Research may provide practical solutions that are generalising rather than brittle, and for which we know the domain in which the solutions should and should not work. We show how, by moving their focus towards Use-Inspired Basic Research, researchers currently focused on basic or applied research respectively, can increase their rate of scientific innovation. We believe that understanding the drivers of scientific innovation and the role of use-inspired basic research can increase the impact of interpretability research towards AI safety and other socially beneficial goals.