An Automated Pipeline for Provably Retrieval-Dependent Benchmark Construction over Dynamic Knowledge
Abstract
Static benchmarks often conflate memorization with reasoning, failing to capture the dynamic nature of world knowledge. We present LiveSearchBench, a scalable pipeline that automatically constructs provably retrieval-dependent benchmarks from knowledge graph differentials. Unlike prior template-based approaches, our method leverages topology-aware synthesis to generate multi-constraint compositional questions with unique answers guaranteed via SPARQL validation. Evaluating 1,000 questions across three difficulty levels on both open-weight and frontier models, we expose a pronounced ``Recency Gap'' on novel facts, particularly for compositional queries. Through Oracle and Wikidata-as-corpus analyses, we disentangle dual failure modes: retrieval coverage limitations for single-hop questions and reasoning failures for compositional queries even with ground-truth evidence. LiveSearchBench provides both a rigorous construction methodology and a continuously updatable testbed for evaluating evidence integration under evolving knowledge.