EngineLab: Evaluating Strategic Generalization Under Rule Shifts
Tianyi Gu ⋅ Lucas Yuan
Abstract
Strategies that work well in one game don't necessarily transfer to another. But, when and why they fail is poorly understood. We introduce EngineLab, a benchmark designed to study this question using chess variants as controlled perturbations of the traditional rule set. In EngineLab, we assign agents with different combinations of strategic weights and compete them in round-robin tournaments, with each feature's importance measured through its Banzhaf power index. Using seven chess variants as rule perturbations, we show that dominant strategies do vary across the games: 6 of 11 features change sign, and the top-ranked feature differs in 5 of 7 variants. An EXP3 bandit learner discovers optimal strategies at rates varying by $10\times$ across variants. A full cross-variant transfer matrix (42 directed pairs, 5 seeds each) shows that warm-starting reliably hurts when the source variant has an inverted objective, while it helps between structurally similar regimes. GPT-4o, when asked to predict feature rankings from rule descriptions, achieves near-random accuracy ($\tau = +0.08$), exhibiting the same overgeneralization from standard-game priors. These results suggest that rule-shifted games offer a useful benchmark for studying strategic transfer.
Chat is not available.
Successful Page Load