Identifying dependent components from multi-domain linear mixtures
Abstract
We study a linear mixing model with dependent latent components, assuming multiple data domains. Most existing models assume that the components are independent or at least uncorrelated, in line with independent component analysis (ICA). Some recent work allows for dependent components, but then makes specific assumptions such as parametric forms of dependencies, multi-view settings, or interventions, or does not recover the individual components. In contrast, we consider a multi-domain setting in which domains differ through domain-specific scalings of the components, while the distribution of the underlying latent components is the same across domains. This approach can model data collected, for example, from different sensors measuring the same process, different laboratories conducting the same experiment, different experimental conditions, or different subjects that might differ in biological or physiological factors. We show that, under sufficient domain variability, latent variables and mixing functions can be identified from second-order statistics alone. We propose the Multi-Domain Covariance Matching (MuDo-CoM) algorithm that generalizes previous methods of joint diagonalization. MuDo-CoM is validated on simulated data and a real-world fMRI dataset.
Lay Summary
In many real-world applications, the data we observe are high-dimensional observations of a few latent underlying factors. For example, functional magnetic resonance imaging (fMRI) measures brain activity in specific regions. In our work, we focus on recovering these latent factors from the observations. Most widely-used methods assume the factors are independent, which is a restriction we want to relax. In these settings, often the distributions of the collected data vary across domains, e.g., due to differences in anatomical and physiological factors, such as age or disease state. We assume we have access to such data from multiple domains, where the only change is in the scale of the latent variables. We prove that these factors can still be uniquely recovered even if they are not independent using only covariance information, a simple statistical property of the data. Consequently, we develop a new method, MuDo-CoM (Multi-Domain Covariance Matching), which extends classical signal separation techniques to handle dependent latent components across multiple domains. Our work broadens the range of situations in which we can recover latent factors in settings where these factors can be correlated. This could improve data analysis methods in areas such as neuroscience, biomedical research, and signal processing.