Multi-Integration of Labels Across Categories for Component Identification in Multi-trial Time Series
Abstract
Many fields collect large-scale temporal data through repeated measurements (`trials’), where each trial is labeled with a set of metadata variables spanning several categories. For example, a trial in a neuroscience study may be linked to a value from category (a): task difficulty, and category (b): animal choice. A critical challenge in time-series analysis is to understand how these labels are encoded within the multi-trial observations, and disentangle the distinct effect of each label entry across categories. Here, we present MILCCI, a novel data-driven method that i) identifies the interpretable components underlying the data, ii) captures cross-trial variability, and iii) integrates label information to understand each category's representation within the data. MILCCI extends a sparse per-trial decomposition that leverages label similarities within each category to enable subtle, label-driven cross-trial adjustments in component compositions and to distinguish the contribution of each category. MILCCI also learns each component’s corresponding temporal trace, which evolves over time within each trial and varies flexibly across trials. We demonstrate MILCCI’s performance through both synthetic and real-world examples, including voting patterns, online page view trends, and neuronal recordings.
Lay Summary
Across many scientific fields, scientists collect repeated measurements over time, each tagged with multiple descriptive labels. For example, brain activity during a task might be recorded over hundreds of repetitions, each tied to both the task outcome and task difficulty. A central question is whether and how each measurement reflects these different labels within its structure, which can be challenging to discern, in particular for measurements that are high-dimensional and vary across repetitions. We developed MILCCI, a method that breaks each measurement into a small set of interpretable components. Each component is tied to a specific experimental factor and has its own structure and temporal pattern. Measurements that share the same value for one factor (e.g., both easy tasks) will share the components associated with that factor, even if they differ in other components. We first demonstrated MILCCI on voting patterns and Wikipedia pageviews, where it discovered components that are aligned with known structure. We then applied it to neural recordings, where it revealed how different aspects of brain activity separately reflect time, task difficulty, and animal's choice. We believe MILCCI can benefit any field where there is a need to understand how multiple experimental factors jointly shape the data.