Skip to yearly menu bar Skip to main content


Poster

DE-COP: Detecting Copyrighted Content in Language Models Training Data

AndrĂ© Duarte · Xuandong Zhao · Arlindo Oliveira · Lei Li


Abstract: *How can we detect if copyrighted content was used in the training process of a language model, considering that the training data is typically undisclosed?* We are motivated by the premise that a language model is likely to identify verbatim excerpts from its training text. We propose DE-COP, a method to determine whether a piece of copyrighted content is included in training. DE-COP's core approach is to probe an LLM with multiple-choice questions, whose options include both verbatim text and their paraphrases. We construct BookTection, a benchmark with excerpts from 165 books published prior and subsequent to a model's training cutoff, along with their paraphrases. Our experiments show that DE-COP outperforms the prior best method by 8.6% in detection accuracy (AUC) on models with logits available. Moreover, DE-COP also achieves an average accuracy of 72% for detecting suspect books on fully black-box models where prior methods give $\approx$ 0\% accuracy. Our code and datasets are available at https://anonymous.4open.science/r/DE-COP-9F1E/

Live content is unavailable. Log in and register to view live content