Scaffolding, Not Stronger Reasoning: Diagnosing Skill-Card Retrieval for Mathematical Problem Solving
Abstract
Skill-card retrieval is often presented as reusable reasoning memory: a solver retrieves compact procedural guidance distilled from prior solutions and conditions on it at test time. We ask whether this mechanism actually strengthens mathematical reasoning, or mainly acts as scaffolding for models that still need help. In a paired diagnostic study with released TRS-style DeepMath skill cards, retrieval improves weaker solvers substantially (+10.4 pp for GPT-4o-mini and +4.2 pp for GPT-5.4), but gives only a small and unreliable gain for GPT-5.5 (+0.6 pp). When exact source cards are removed, these gains do not persist, and the GPT-5.5 condition falls 2.6 pp below Direct. For the strongest solver in our study, lower-ranked and random cards shift paired flips toward harm, and manual audits show that many failures arise from partially applicable skills that introduce caveats, convention shifts, wrong limiting regimes, or symbolic drift. These findings suggest that skill-card retrieval is best viewed as useful scaffolding for weaker solvers rather than as a uniformly reliable improvement as solver capability increases.