A Comparative Study on Abstractive and Extractive Approaches In Summarization of European Legislation Documents

4Citations
Citations of this article
46Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Extracting the most important part of legislation documents has great business value because the texts are usually very long and hard to understand. The aim of this article is to evaluate different algorithms for text summarization on EU legislation documents. The content contains domain-specific words. We collected a text summarization dataset of EU legal documents consisting of 1563 documents, in which the mean length of summaries is 424 words. Experiments were conducted with different algorithms using the new dataset. A simple extractive algorithm was selected as a baseline. Advanced extractive algorithms, which use encoders show better results than baseline. The best result measured by ROUGE scores was achieved by a fine-tuned abstractive T5 model, which was adapted to work with long texts.

Cite

CITATION STYLE

APA

Zmiycharov, V., Chechev, M., Lazarova, G., Tsonkov, T., & Koychev, I. (2021). A Comparative Study on Abstractive and Extractive Approaches In Summarization of European Legislation Documents. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 1645–1651). Incoma Ltd. https://doi.org/10.26615/978-954-452-072-4_184

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free