Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics

3Citations
Citations of this article
42Readers
Mendeley users who have this article in their library.

Abstract

Question answering-based summarization evaluation metrics must automatically determine whether the QA model's prediction is correct or not, a task known as answer verification. In this work, we benchmark the lexical answer verification methods which have been used by current QA-based metrics as well as two more sophisticated text comparison methods, BERTScore and LERC. We find that LERC out-performs the other methods in some settings while remaining statistically indistinguishable from lexical overlap in others. However, our experiments reveal that improved verification performance does not necessarily translate to overall QA-based metric quality: In some scenarios, using a worse verification method - or using none at all - has comparable performance to using the best verification method, a result that we attribute to properties of the datasets.

Cite

CITATION STYLE

APA

Deutsch, D., & Roth, D. (2022). Benchmarking Answer Verification Methods for Question Answering-Based Summarization Evaluation Metrics. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 3759–3765). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2022.findings-acl.296

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free