NOVA-Test: Auditing LLM-Generated Research Hypotheses with Structural Analogy, Prior-Work, and Falsifiability Tests
Abstract
We study LLM research ideation as a hypothesis-testing problem, not a prompting tutorial. The unit under test is a generated research hypothesis together with its structural justification, nearest-prior relation, and proposed falsifier. We repurpose NOVA-Analog as NOVA-Test, a lightweight rejection protocol with three gates: source-target structural alignment, nearest-prior novelty audit, and explicit falsification planning. Across 12 ML/CS bottlenecks and three generation conditions, NOVA-Test improves judged novelty (3.88 vs. 3.00 one-shot and 2.83 literature-search), technical specificity (4.54 vs. 3.21 and 3.00), and falsifiability (4.79 vs. 4.46 and 4.21). Bonferroni-corrected Wilcoxon tests remain significant for all planned comparisons except feasibility versus literature-search. A 488-work novelty audit labels 10/12 NOVA-Test hypotheses plausibly novel versus 3/12 literature-search hypotheses (Fisher one-sided p=0.006). The contribution is not autonomous discovery; it is a reproducible hypothesis-audit scaffold for LLM-related testing problems.