Tacit Verification Apprenticeship for Self-Evolving Scientific Agents
Abstract
Self-evolving scientific agents will not be reliable merely because they can call proof assistants, validators, model checkers, or executable workflows. The hard part is a missing layer of tacit verification craft: choosing definitions, preserving informal intent, splitting proof obligations, navigating libraries, repairing failed scripts, and knowing when a formally accepted artifact is scientifically misaligned. We introduce Tacit Verification Apprenticeship (TVA), an architecture and evaluation agenda in which agents learn this craft from proof-state trajectories, failed attempts, expert corrections, semantic-alignment checks, and domain validation traces. We instantiate the idea with a typed memory schema, utility model, failure taxonomy, benchmark design, and three reproducible corpus probes over 1,406 curated records plus an executable semantic-alignment demo. The result is a practical blueprint for scientific agents that improve across cycles while remaining checkable, auditable, and semantically faithful.