EO-Agents: A Three-Agent LLM Pipeline for Earth Observation Hypothesis Generation
Abstract
Large language models have recently been explored for scientific hypothesis generation, but most prior approaches rely on unstructured literature corpora and free-form textual claims. We present a pipeline for Earth observation that grounds hypothesis generation directly in the NASA Earth Observation Knowledge Graph. A heterogeneous graph neural network trained on historical dataset co-usage relations ranks candidate dataset pairings, while a three-agent large language model pipeline filters, generates, and evaluates structured research hypotheses. Applied to 1,475 NASA datasets, the system produces 160 hypotheses spanning ecohydrology, glaciology, aerosol–cloud interactions, vegetation phenology, and methodological bias correction. Model-predicted novel dataset pairings are rated nearly as plausible as held-out real co-usages from the literature, suggesting that the pipeline identifies scientifically coherent yet previously unexplored combinations. A 2 × 2 × 2 factorial experiment across GPT-5.2 and Claude Sonnet 4.6 further shows that relative hypothesis rankings remain stable across configurations, whereas absolute evaluation scores depend strongly on judge identity, highlighting limitations of single-judge LLM-based evaluation.