Skip a Layer or Loop It? Learning Program-of-Layers in LLMs
Abstract
Large language models (LLMs) perform inference by following a fixed depth and order, non-recurrent execution of all layers. We reveal the wide existence of training-free, flexible, dynamic program-of-layers (PoLar), where pretrained layers can be packed as modules and then skipped or looped to form a customized program for each input. For most inputs, substantially shorter program executions can achieve the same or better accuracy, while incorrect predictions of the original LLM can be corrected by alternative programs with fewer layers. These observations indicate that inference admits multiple valid latent computations beyond the standard forward pass. To efficiently achieve PoLar in practice, we propose a lightweight PoLar prediction network, which learns to generate execution programs that dynamically skip or repeat pretrained layers for each input. Experiments on mathematical reasoning benchmarks demonstrate that PoLar consistently improves accuracy over standard inference and prior dynamic-depth methods, often while executing fewer layers, and that these gains persist under out-of-distribution evaluation. Our results suggest that fixed-depth execution captures only a narrow subset of an LLM’s latent reasoning capacity.
Lay Summary
Large language models usually answer every question by running through the same fixed sequence of layers, even though some questions are easy and others are much harder. This paper studies whether a model can use its existing layers more flexibly at inference time. We show that, for many inputs, a model can skip some layers or reuse some layers and still produce the correct answer, sometimes with better accuracy and less computation. Based on this observation, we propose PoLar, a lightweight method that predicts a customized layer execution plan for each input without changing the original model weights. Experiments on math reasoning and out-of-distribution benchmarks show that this adaptive use of layers can improve accuracy and efficiency. Our results suggest that pretrained language models contain useful computational paths beyond the standard fixed forward pass.