POLYOMINOGEN: A Controlled Testbed for Understanding Memorization and Compositional Generalization in Conditional Diffusion Models
Abstract
Deep generative models can produce samples that appear novel, yet in natural-image domains it is often unclear whether such samples reflect memorization, interpolation, or systematic generalization beyond the observed training support. We introduce PolyominoGen, a controlled testbed for studying memorization and compositional generalization in conditional diffusion models. Each image is rendered from an exact symbolic tuple specifying shape, color, orientation, and position, enabling train--test splits in which specific attribute combinations or geometric transformations are withheld by construction. This known-support setting allows generated samples to be evaluated using rule-based validity, conditional tuple accuracy, nearest-neighbor proximity, and novel-valid held-out generation metrics. In a preliminary U-Net DDPM pilot, we observe that high in-support validity does not imply held-out compositional accuracy, and that few-shot adaptation from a pretrained model can outperform training from scratch on held-out tuples. PolyominoGen provides a lightweight diagnostic framework for probing when generative models memorize, when they generalize, and when they fail under controlled distribution shift.