Pushing the Limitation of High-order Interaction in Language Models Using Tensor Networks
Zhan Su ⋅ Fengran Mo ⋅ Guojun Liang ⋅ Prayag Tiwari
Abstract
Neural network language models (NNLMs) have drawn much attention in AI research. The current paradigm for language models is mainly based on transformer-based structure or recurrent neural networks (RNN), where each word is represented by a vector representation $X\in \mathbb{R}^d$. We posit that this limits the capacity of the language model, which shows limitations for datasets where the representation is based on \textit{high-order interactions}. To address this, we investigate a novel language model structure based on high-order tensor representation $\mathcal{X}\in\mathbb{R}^{d_1\times d_2\times...\times d_n}$, followed by a tensor train decomposition. We term this `Tensor Train Language Model' (TTLM). TTLM represents sentences in an exponential space constructed by the tensor product, but it computes the probabilities of sentences in a low-dimensional fashion. We evaluate the TTLM on high-order molecule tasks and language model tasks. Experimental results demonstrate that the proposed variants of TTLM (i.e., TTLM-Large and TTLM-Tiny) outperform the Recurrent Neural Networks (RNNs) and vanilla transform-based structure.
Chat is not available.
Successful Page Load