Paper #32: Can Large Language Models Match Human Diversity in Educational Content Generation?
Abstract
Large language models (LLMs) are known to lack diversity in open-ended text generation, a limitation that is especially consequential in education, where the breadth of instructional moves shapes the kinds of thinking and disciplinary work students are asked to do. Across two educational content generation tasks, math items and essay feedback, we evaluate diversity using lexical, semantic, and domain-grounded metrics. For math items, we measure diversity in instructional objectives and application contexts; for essay feedback, we measure diversity in content focus and discourse form. We evaluate frontier and open-weight models using ordinary task prompts, diversity-oriented prompting interventions, and post-training. Across both tasks, baseline model generations exhibit on average 67–81% more intra-model lexical similarity than human-written reference distributions. Prompting improves diversity mainly through surface variation. Post-training yields inconsistent gains and can reduce output acceptability. These results suggest that current efforts to mitigate mode collapse are insufficient for open-ended educational generation, and motivate new training and data collection strategies to support pedagogical diversity.