Abstract
Existing models for extractive summarization are usually trained from scratch with a cross-entropy loss, which does not explicitly capture the global context at the document level. In this paper, we aim to improve this task by introducing three auxiliary pre-training tasks that learn to capture the document-level context in a self-supervised fashion. Experiments on the widely-used CNN/DM dataset validate the effectiveness of the proposed auxiliary tasks. Furthermore, we show that after pretraining, a clean model with simple building blocks is able to outperform previous state-of-the-art that are carefully designed.
Cite
CITATION STYLE
Wang, H., Wang, X., Xiong, W., Yu, M., Guo, X., Chang, S., & Wang, W. Y. (2020). Self-supervised learning for contextualized extractive summarization. In ACL 2019 - 57th Annual Meeting of the Association for Computational Linguistics, Proceedings of the Conference (pp. 2221–2227). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/p19-1214
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.