Does My Embedding Reflect That \(A = B\)? Evaluating Mathematical Equivalence in Embedding Models
Abstract
In this paper, we investigate whether current text embedding models can capture mathematical equivalence. To do this, we introduce the \emph{Mathematically Equivalent but Lexically Different Pairs (MELD) Dataset}, a collection of paired statements that are expressed in very different language but are mathematically equivalent. We show that the current state-of-the-art embedding models struggle with this task. Motivated by this, we propose a contrastive approach to learning embeddings of mathematical text that focuses on aligning different representations. We specifically train models to align informal statements with different formalizations. Our experiments demonstrate that this leads to improvements not only on informal-formal retrieval tasks but also on MELD, which only contains natural language statements.