Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex
Abstract
Task-optimized neural networks are the leading in-silico models of sensory cortex, yet the field lacks a unified understanding of which modeling choices drive improved brain alignment. Prior NeuroAI work is fragmented across datasets and modalities, making it difficult to determine robust scaling trends. Here, we systematically investigate the scaling laws of model-to-brain alignment across 8 neural datasets (spanning electrophysiology, fMRI, EEG, and MEG) and over 600 models with diverse architectures and pretraining configurations. We report three scaling trends: (1) Pretraining saturation: Alignment improves with pretraining compute and data scale but saturates across all recording modalities. (2) Complementary fine-tuning: Hybrid task & neural data optimization yields consistent improvements in alignment that generalize across datasets and modalities. (3) Mapping scaling: Increasing the number of neural samples to fit model-to-brain mappings yields log-linear gains with the largest impact on alignment. Finally, we propose a novel subject-shared cross-attention mapping which drastically reduces parameter count and improves alignment. Taken together, these results establish multimodal scaling laws that guide resource allocation for next-generation brain models.
Lay Summary
AI vision models are among our best tools for predicting how neurons respond to images, yet it remains unclear what makes one model more brain-like than another. Should we train bigger networks on more images, expose them to brain recordings, or improve how we translate model activity into neural signals? Past studies disagree because they use different datasets and pipelines. We ran a unified experiment across 600+ vision models and eight neural datasets spanning monkey electrodes and human fMRI, EEG, and MEG. For each model, we measured brain alignment while varying three ingredients: training scale, fine-tuning on neural recordings, and the amount of paired image-and-brain data used for mapping. Bigger training runs helped but eventually saturated; neural training added a small consistent boost; and more paired brain data produced the largest, most reliable gains. We also introduce a compact mapping method that shares information across people while using far fewer fitted parameters. These results turn vague intuitions about scaling brain models into a concrete map of where extra compute or data pays off most. This guides researchers in investing scarce neuroscience datasets where they will most improve our digital models of vision.