Transferable Lesion-Supervised Speech Representations for Post-Stroke Modelling
Abstract
Characterising post-stroke brain injury typically relies on structural neuroimaging, which is costly, infrastructure-dependent, and poorly suited to repeated or large-scale monitoring; this also creates a data-efficiency problem, because supervised targets in clinical cohorts are often sparsely observed across modalities, limiting the amount of paired speech and clinical data available for downstream modelling. We investigate a transfer learning approach in which representations learned through lesion inference from speech are reused for downstream structural and clinical phenotypes absent from the original training signal. These lesion-supervised representations are derived from a multi-head speech-to-lesion (S2L) model as patient-level aggregations of out-of-fold predicted lesion probabilities, stacked learner outputs, and uncertainty estimates across speech samples. These S2L representations are benchmarked against clinically interpretable speech features, Whisper embeddings, clinical covariates, and demographic baselines under nested cross-validation. S2L representations are the strongest direct predictors of focal lesion bounding-box extent ((R^2=0.241), Pearson (r=0.503)), despite spatial extent being unseen during S2L training, while high-dimensional Whisper encoder embeddings perform best directly for cognitive outcomes ((R^2=0.486), (r=0.702)) with further improvement after residual adaptation using S2L representations ((R^2=0.501)). These findings suggest that representation transfer can partially offset sparse clinical annotations, offering a scalable route for modelling post-stroke brain--behaviour relationships.