CJEUSumm: A Multilingual Dataset for the (Long-Document) Summarization of EU Court Decisions Based on Official Press Releases
Piotr Rataj ⋅ Bianca Steffes ⋅ Christoph Sorge ⋅ Kristof Meding
Abstract
We compile a high-quality multilingual dataset for abstractive legal text summarization. The dataset pairs decisions of the Court of Justice of the European Union (CJEU) with official press releases. The corpus comprises 26 218 decision–summary pairs with relatively long input ($\sim$10k tokens) and output ($\sim$900 tokens), and with extensive parallel coverage in up to 23 languages. Using a subset of the data containing English, German and Polish data points, we benchmark GPT-4.1, GPT-4.1-nano, GPT-5 and GPT-5-nano in a zero- and one-shot setting and find comparatively strong performance on standard automated metrics. An anecdotal manual inspection does not reveal substantial faithfulness issues.
Chat is not available.
Successful Page Load