GeoRecon: Graph-Level Representation Learning for 3D Molecules via Reconstruction-Based Pretraining
Abstract
Pretraining–-finetuning has driven major advances in natural language processing and vision through objectives such as masked language modeling and next-token prediction. In molecular representation learning, however, pretraining tasks remain largely restricted to node-level denoising, which captures local atomic environments but provides no explicit training signal for the global molecular structure needed by graph-level property prediction tasks such as energy estimation and molecular regression. To address this gap, we introduce GeoRecon, a graph-level pretraining framework that shifts the reconstruction target from individual atoms to the molecule as an integrated whole. During pretraining, GeoRecon learns a graph representation that conditions geometry reconstruction and induces smoother, more transferable latent spaces. This encourages coherent global structural features beyond isolated atomic details while remaining fully self-supervised and 3D-only. Across QM9, MD17, MD22, and appendix benchmarks, GeoRecon consistently improves over its direct coordinate-denoising baseline and remains competitive with broader molecular pretraining baselines, supporting graph-level reconstruction as a simple and effective complement to node-level denoising.