Mining Useful General Data for Low-Resource Domain Adaptation
Abstract
Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data is limited, there exists a vast amount of general-domain data that shares similar question–answer formats and reasoning patterns with domain tasks. This observation raises an important question: can useful general-domain data be mined to improve low-resource domain adaptation? Our initial findings show that general-domain chain-of-thought data contains useful auxiliary signals for domain adaptation, even without careful selection. This observation motivates a new paradigm for domain adaptation beyond exclusive reliance on domain-specific data. To systematically identify the most beneficial general-domain samples, we propose NTK-Selector, motivated by the Neural Tangent Kernel’s ability to capture alignment in training dynamics. Since directly applying NTK to pretrained LLMs is impractical, we introduce a Jacobian-free NTK approximation and empirically demonstrate stable NTK-like behavior during fine-tuning. Extensive experiments across medical, financial, legal, and psychological domains demonstrate that NTK-Selector consistently outperforms domain-only fine-tuning and existing data selection baselines. In particular, NTK-Selector achieves gains of +8.7 and +5.1 points on Llama3-8B-Instruct and Qwen3-8B, respectively, compared to only +0.8 and +0.9 points from domain-only fine-tuning.
Lay Summary
Large language models are often much harder to use in specialized areas such as medicine, law, finance, or mental health, because these fields usually have very little training data. However, there is a huge amount of general-purpose data on the internet and in public datasets. Some of this data may still be helpful, because it can contain similar question-answer styles or reasoning steps. In this paper, we ask a simple question: can we find the helpful parts of general data and use them to improve performance in data-scarce domains? To do this, we propose a method called NTK-Selector, which tries to pick the general examples that are most useful for learning the target task. Our results show that this approach works better than training only on the small amount of domain data, and better than several existing data selection methods. In practical terms, this means we may be able to build stronger domain-specific AI systems while relying less on expensive expert-labeled data.