Active Timepoint Selection for Learning Measure-Valued Trajectories
Abstract
Inferring continuous probability paths from sparse snapshots is a fundamental challenge in domains like single-cell biology, where high-fidelity data acquisition is often destructive and constrained by prohibitive sequencing costs. This motivates the need for active learning strategies to strategically select optimal measurement times. However, designing active learning policies for this setting remains an open problem: the target objects reside on the infinite dimensional Wasserstein space where standard Euclidean metrics are ill-defined, and current interpolation methods lack epistemic uncertainty quantification. We introduce a framework which extends active experimentation to the space of measures. By leveraging Linearized Optimal Transport (LOT), we map distributional snapshots into a tangent space amenable to Gaussian Process modeling, allowing us to construct a tractable probabilistic surrogate for the underlying probability path. This yields an acquisition policy that iteratively selects measurement times to minimize uncertainty. Empirical results demonstrate that our strategy outperforms uncertainty-agnostic baselines on both synthetic and real-world datasets.
Lay Summary
Many scientific experiments try to understand how a whole population changes over time, such as cells transforming from one type into another. In single-cell biology, each measurement can be expensive and destructive, so scientists often cannot measure every possible time point. This raises a practical question: when should we take the next measurement to learn the most about the entire process? We study this question for data where each time point is not a single number, but a distribution, such as a cloud of thousands of cells. Our method turns these distributions into a simpler geometric representation, uses a probabilistic model to estimate both the likely trajectory and where it is still uncertain, and then chooses the next time point where a new measurement should be most useful. It also adapts to processes that change unevenly, with long quiet periods followed by rapid branching events. In experiments on simulated data and real single-cell reprogramming data, our approach reconstructs the changing distributions more accurately than choosing time points uniformly or at random, especially when only a modest number of measurements is available. This could help researchers design more informative and cost-effective time-course experiments.