Evaluating Surprise Adequacy for Question Answering

Seah Kim; Shin Yoo

Conference ProceedingsOPEN ACCESS

Evaluating Surprise Adequacy for Question Answering

Proceedings - 2020 IEEE/ACM 42nd International Conference on Software Engineering Workshops, ICSEW 2020 (2020) 197-202

DOI: 10.1145/3387940.3391465

14Citations

15Readers

Get full text

Abstract

With the wide and rapid adoption of Deep Neural Networks (DNNs) in various domains, an urgent need to validate their behaviour has risen, resulting in various test adequacy metrics for DNNs. One of the metrics, Surprise Adequacy (SA), aims to measure how surprising a new input is based on the similarity to the data used for training. While SA has been evaluated to be effective for image classifiers based on Convolutional Neural Networks (CNNs), it has not been studied for the Natural Language Processing (NLP) domain. This paper applies SA to NLP, in particular to the question answering task: the aim is to investigate whether SA correlates well with the correctness of answers. An empirical evaluation using the widely used Stanford Question Answering Dataset (SQuAD) shows that SA can work well as a test adequacy metric for the question answering task.

Author supplied keywords

Cite

CITATION STYLE

APA

Kim, S., & Yoo, S. (2020). Evaluating Surprise Adequacy for Question Answering. In Proceedings - 2020 IEEE/ACM 42nd International Conference on Software Engineering Workshops, ICSEW 2020 (pp. 197–202). Association for Computing Machinery, Inc. https://doi.org/10.1145/3387940.3391465

Evaluating Surprise Adequacy for Question Answering

Abstract

Author supplied keywords

Cite

Register to see more suggestions