LLMs Can Learn the Language of the Microbiome
Abstract
We explore the application of large language models (LLMs) to microbiome data, a domain that remains underexplored despite the rise of self-supervised learning in biology. We introduce Atlas, a large-scale pretraining dataset comprising over 539,000 data points from MGnify, spanning multiple DNA sequencing modalities including amplicon, assembly, and whole-metagenome data. Using Atlas, we train the Waypoint family of models, GPT-style causal language models trained to understand microbiomes. To enable standardized evaluation, we present Compass, a benchmark of eight downstream microbiome prediction tasks. We show that our pretrained Waypoint models outperform classical methods and prior foundation models, with gains driven by both dataset scale and representation choices. Our results establish pretrained LLMs as a strong and practical approach for microbiome prediction tasks.