Cross-Subject Modeling for Widefield Calcium Imaging via Atlas-Aligned Spatiotemporal Tokenization
Abstract
Large-scale, multi-subject widefield calcium imaging provides unprecedented access to brain-wide cortical dynamics. However, the high dimensionality, complex spatiotemporal structure, and substantial task-irrelevant activity in widefield recordings have largely restricted modeling efforts to single-session analyses, limiting scalability and generalization. While multi-subject pretrained models have been explored for some neural modalities, multi-subject models for widefield calcium imaging have not yet been demonstrated; further, subject-invariant zero-shot behavior decoding remains elusive for multi-subject models across neural modalities more broadly. As a first step toward foundation modeling of widefield data, we introduce WiCAT, a multi-subject model that leverages self-supervised pretraining to both outperform single-session models and enable zero-shot behavior decoding on unseen subjects. WiCAT introduces an atlas-grounded tokenization scheme without session-specific components and learns globally shared spatiotemporal representations. Across multiple widefield datasets, the pretrained model supports lightweight downstream decoding, transfers across subjects, tasks, and datasets, and outperforms baseline models. Notably, the model also achieves robust zero-shot continuous behavior decoding and left-out brain region reconstruction on unseen subjects.
Lay Summary
Modern brain imaging can record activity across large parts of the brain while subjects perform various behaviors, giving scientists a powerful way to study how brain-wide activity gives rise to actions. But building models that generalize from one subject to another remains a major challenge in neuroscience. Brain imaging data are high-dimensional videos, contain complex patterns that unfold across both space and time, and vary across subjects, recording sessions, and experiments. As a result, most existing methods train separate models for each recording session or subject, making it difficult to reuse what a model has learned when studying a new subject. We present WiCAT, a model designed to learn shared brain-wide representations from large widefield calcium imaging datasets across subjects. WiCAT aligns brain activity videos to a common brain atlas, divides these videos into spatial and temporal patches, and trains a transformer model by masking large parts of brain images and reconstructing them from the remaining context. Because WiCAT is trained without subject- or session-specific identifiers, it is encouraged to learn general structure in brain-wide activity rather than memorizing data from individual recording sessions. We show that WiCAT can decode continuous behavior – including movement of various body parts – in completely unseen subjects without any retraining on those subjects, a form of zero-shot transfer that has remained elusive in neuroscience because neural recordings often differ substantially across subjects and sessions. WiCAT also reconstructs activity in left-out brain regions of unseen subjects using the rest of the brain as context, showing that it learns transferable relationships across brain areas. Overall, WiCAT provides a first step toward reusable foundation modeling for widefield brain imaging, where one model can learn from large neural datasets and transfer to new subjects with little or no additional training.