Persistent Homology for Distribution Drift Detection in LLM Embedding Streams
Lei Yang
Abstract
Deployed language models experience distribution drift in their input streams, degrading performance silently. Existing drift detectors rely on low-order statistics---means, covariances, or kernel embeddings---and may underperform when drift alters the geometric structure of the embedding space without large changes in these quantities. We present the first systematic evaluation of persistent homology for drift detection in LLM embedding streams. We design three centroid-controlled drift scenarios---subtopic reweighting (approximate centroid preservation), style perturbation, and geometric reorganization (exact centroid preservation)---that reduce the advantage of conventional detectors, and evaluate 8 topological features against 6 classical baselines across 2 datasets, 2 embedding models, and 5 random seeds. Our key finding is that Wasserstein distance on $H_0$ persistence diagrams achieves $\text{AUC} = 0.858$ on geometric reorganization drift, significantly outperforming the best classical baseline ($\text{MMD-RBF}$, $\text{AUC} = 0.782$; paired permutation test $p < 0.001$). However, the advantage is scenario-specific: classical methods remain superior on subtopic reweighting drift, and all methods show only moderate power on style perturbation. We provide comprehensive ablations over window size, subsample size, and PCA dimensionality, and show that opological methods run at $37\text{--}45$ ms per window under the tested settings. Our results establish when topological drift detection adds complementary value and when classical methods suffice.
Chat is not available.
Successful Page Load