SymboLLM-FE: LLM Accelerated Symbolic Regression for Automated Feature Engineering
Zijian Cheng ⋅ 贾 子怡 ⋅ Zhi Zhou ⋅ Yu-Feng Li ⋅ Lan-Zhe Guo
Abstract
Automated Feature Engineering (AutoFE) can algorithmically generate features and excel in replacing manual feature construction to reduce human effort and improve scalability. However, traditional AutoFE suffer from high computational costs and poor interpretability due to combinatorial explosion in the search space, while large language models (LLM)-based AutoFE faces reliability issues such as high time cost and hallucination bias. In this paper, we combine **symbo**lic regression with L**LM**s for **f**eature **e**ngineering (SymboLM-FE) to solve these challenges. We extract mathematically expressive rules strongly correlated with the target via symbolic regression, then refine them using LLMs with rich prior knowledge, forming a closed-loop AutoFE framework. Empirical results on six real-world datasets and four Kaggle competitions demonstrate that SymboLM-FE outperforms existing AutoFE by a large margin. SymboLM-FE also mitigates combinatorial explosion $\mathcal{O}(n^2)$ via expanding-sliding search and enhances interpretability via statistical prior-grounded closed-loop LLM refinement.
Chat is not available.
Successful Page Load