Auditing Generative Graph Foundation Models for Connectomics: A Score × Predictor × Sampler Decomposition on Real Sparse Directed Connectomes
Abstract
Generative graph foundation models are now deployed across the life sciences as in-silico instruments for biological wiring: pre-fill simulation-ready connectomes on partially imaged volumes, hypothesize wiring at scales no current electron microscopy volume reaches, and infer specific wiring rules. The dominant deployment recipe pairs a learned per-edge score from a generative graph backbone (DiGress, DeFoG, GDSS, SPECTRE, DIRECTO) with a non-parametric Sinkhorn–Knopp degree rescaling step that enforces target in-out degree marginals on each new graph. We audit this pipeline on two real sparse directed connectomes spanning a tenfold density range — MICrONS Phase-1 layer 2/3 mouse visual cortex and the FlyWire central plus right hemisphere — and report a score × predictor × sampler decomposition with cluster-bootstrap confidence intervals and Benjamini–Hochberg false-discovery-rate control on a four-axis evaluation suite. The four axes are deliberate: three-hop reachability and triadic motif distance probe marginal-summary biological structure; PolyGraph Discrepancy is a learned-classifier axis that approximates the joint distribution; persistent-homology bottleneck distance for one-cycles is a topology-fidelity axis. The Pareto verdict is metric-specific. On marginal-summary axes the front contains exactly two configurations, both structural nulls that use the test-graph degree sequence; every feature-based deployable predictor we tested is strictly Pareto-dominated. On the learned-classifier axis the deployable RandomForest plus Sinkhorn pipeline beats the structural null; on the persistent-homology axis it ties the structural null. A faithful in-substrate DiGress-style baseline composed with Sinkhorn becomes a new motif-axis Pareto-front member. We further probe the structured-projection question that is load-bearing for any biological deployment of a graph foundation model: how much of the trained backbone's score actually survives the Sinkhorn projection. Across input scores, distance-kernel and pairwise-MLP scores retain almost none of their top-edge rank under the projection, but a learned discrete flow score retains a non-negligible fraction — the only score we tested whose Sinkhorn-survival is meaningfully above noise. The implication for foundation models for the life sciences is that score-side capacity gains have low expected return under this particular deployment composition unless the new score carries degree-marginal-aware joint structure that the projection partially preserves; a learned discrete flow is the only family we tested that does. We also derive the Decelle sparse stochastic-block-model detectability threshold for both substrates and confirm that the joint axis is structurally undersampled at realistic node counts and densities, which explains why marginal-summary axes alone misrepresent the inference difficulty for biological wiring. The implications for biology and for graph foundation models in the life sciences are concrete. Published evaluations on sparse biological substrates that report only marginal-summary axes substantially overstate the contribution of the trained backbone above the configuration-model structural null. The reporting standard we propose — every feature-based headline reported alongside the configuration-model upper bound, a learned-classifier axis, and a topology-fidelity axis, with the score × predictor × sampler decomposition explicitly listed — makes the deployable contribution auditable. We do not retrain a multi-state-of-the-art sweep at scale; the in-substrate DiGress-analog reported here is the closest direct biological baseline. Two GPU-free deployable recipes are released for the life-sciences-deployment regime.