Paper #26: Do Semantic Distance Tests Actually Predict Creativity in LLMs?
Abstract
Automated semantic distance tests—which prompt a model to produce a set of words, and score the average embedding distance between them—are increasingly used to measure the "creativity" of large language models. However, the validity of semantic distance tests as predictors of machine creativity has not yet been established, and these tests already have limited validity as predictors of human creativity. To address this problem, we conduct the first systematic study evaluating the effectiveness of semantic distance tests in predicting creative achievement across three constructs: creative writing, divergent thinking, and scientific ideation. We score each test on two criteria (validity and specificity), and derive a theoretical limit for the maximum attainable specificity and validity a test can achieve. We find that: (1) Test effectiveness varies significantly by construct, and no single test predicts all constructs well. (2) None of the tests is a good predictor of scientific ideation ability. (3) Existing tests are far below the theoretical limits, indicating meaningful room for the design of improved tests moving forward. Our findings provide clear practical takeaways and directions for future work and suggest that novel tests are needed to reliably predict scientific ideation ability.