Rectified LpJEPA: Joint-Embedding Predictive Architectures with Sparse and Maximum-Entropy Representations
Yilun Kuang ⋅ Yash Dagade ⋅ Tim G. J. Rudner ⋅ Randall Balestriero ⋅ Yann LeCun
Abstract
Joint-Embedding Predictive Architectures (JEPA) learn view-invariant representations and admit projection-based distribution matching for collapse prevention. Existing approaches regularize representations towards isotropic Gaussian distributions, but inherently favor dense representations and fail to capture the key property of sparsity observed in efficient representations. We introduce Rectified Distribution Matching Regularization (RDMReg), a sliced two-sample distribution-matching loss that aligns representations to a Rectified Generalized Gaussian (RGG) distribution. RGG enables explicit control over expected $\ell_0$ norm through rectification, while its continuous truncated component admits a maximum-entropy characterization under expected $\ell_p$ norm and support constraints. Equipping JEPAs with RDMReg yields Rectified LpJEPA, which strictly generalizes prior Gaussian-based JEPAs. Empirically, Rectified LpJEPA learns sparse, non-negative representations with favorable sparsity--performance trade-offs and competitive downstream performance on image classification benchmarks, showing that RDMReg can enforce sparsity while preserving task-relevant information.
Lay Summary
Modern AI systems learn internal representations of data such as images, but these representations are often dense and hard to interpret. We introduce a method for training self-supervised models to learn sparse representations, where only a small number of units are active for each input, while still preserving useful information. On image classification benchmarks, our method produces sparse, non-negative representations with competitive performance, suggesting that AI systems can be encouraged to use more compact and structured internal descriptions without losing task-relevant information.
Successful Page Load