HOLMS: Alternative Summary Evaluation with Large Language Models

13Citations
Citations of this article
76Readers
Mendeley users who have this article in their library.

Abstract

Efficient document summarization requires evaluation measures that can not only rank a set of systems based on an average score, but also highlight which individual summary is better than another. However, despite the very active research on summarization approaches, few works have proposed new evaluation measures in the recent years. The standard measures relied upon for the development of summarization systems are most often ROUGE and BLEU which, despite being efficient in overall system ranking, remain lexical in nature and have a limited potential when it comes to training neural networks. In this paper, we present a new hybrid evaluation measure for summarization, called HOLMS, that combines both language models pre-trained on large corpora and lexical similarity measures. Through several experiments, we show that HOLMS outperforms ROUGE and BLEU substantially in its correlation with human judgments on several extractive summarization datasets for both linguistic quality and pyramid scores.

Cite

CITATION STYLE

APA

Mrabet, Y., & Demner-Fushman, D. (2020). HOLMS: Alternative Summary Evaluation with Large Language Models. In COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Conference (pp. 5679–5688). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.coling-main.498

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free