Beyond Direct Gains: Matched Controls for Evaluating Concept Scaffolds
Jiangshan He ⋅ Song Dai ⋅ Xiaolong Qiao ⋅ Jungang Li ⋅ Yibo Yan ⋅ Xuming Hu
Abstract
Reasoning scaffolds such as concepts and hints are often evaluated by asking whether a helpful scaffold improves accuracy. This direct comparison is under-identified: a gain may come from the scaffold's semantic content, from the presence of a scaffold-shaped prompt block, or from sensitivity to plausible but inapplicable auxiliary information. We propose a matched-control protocol for concept scaffolds that compares no-scaffold, format-only, helpful, misleading, and wrong-fact conditions. We evaluate it on 2,000 undergraduate mathematics problems with three randomized versions per problem, which lets us distinguish average instance-level gains from stable gains across variants of the same reasoning objective. In this nine-setting evaluation, helpful concepts improve over format-only controls for most settings under stable accuracy, but the average gain is modest ($G_{\mathrm{helpful}}=+1.22$ points). We interpret this as a diagnostic helpful-over-format contrast rather than a pure estimate of semantic effect.
Chat is not available.
Successful Page Load