From Noisy Oracles to Useful Constraints: LLM-Guided Constraint Selection for Synthetic Tabular Data
Tejumade Afonja ⋅ Joscha Cüppers ⋅ Mario Fritz
Abstract
Structured foundation models increasingly depend on large corpora of high-quality tabular data for pretraining and evaluation. Yet standard deep generative models routinely violate domain constraints, producing structurally implausible samples that are unsuitable for high-stakes downstream tasks. Constraint-aware generation addresses this but require constraints to be specified upfront, creating a costly annotation bottleneck. We propose EVS (Extract, Verify, Select), which removes this bottleneck by treating large language models as scalable but noisy proxies for domain knowledge. EVS extracts candidate constraints from column metadata, verifies them against the training data, and selects a utility-preserving subset via binary search in $\mathcal{O}(\log K)$ model evaluations without any ground-truth constraint labels. Across three datasets and four LLMs, EVS reduces constraint violation rates to near zero while matching the downstream utility of unconstrained generation, and approaches the performance of expert-annotated baselines when the LLM recovers high-quality constraints.
Chat is not available.
Successful Page Load