$\texttt{LAUGHS}$: An LLM-compatible Molecular String Representation
Gunwook Nam ⋅ Gaeun Sun ⋅ Yousung Jung
Abstract
Large language models (LLMs) are increasingly applied to chemistry, yet their performance depends strongly on how molecules are represented as text. IUPAC names become syntactically unwieldy for complex structures, while graph-serialized strings disperse chemically meaningful moieties across the sequence. Here, we present $\texttt{LAUGHS}$, an LLM-compatible molecular string representation that decomposes a molecule into named moieties, hierarchically organizes them into a tree structure, and linearizes the result into a natural-language-like string. Tokenization analysis reveals that LAUGHS units align near-perfectly with tokenizer spans, suggesting strong compatibility with LLMs. On the property explanation task, LAUGHS matches IUPAC-level performance across all metrics; on site-specific editing, it substantially outperforms all baselines with a 91.4\% exact match rate among valid outputs. Together, our results suggest that semantic mismatch between molecular representations and natural language syntax is a key bottleneck for LLMs in chemistry, and that LAUGHS offers an effective way to address it.
Chat is not available.
Successful Page Load