Abstract
We propose a new dataset for evaluating question answering models with respect to their capacity to reason about beliefs. Our tasks are inspired by theory-of-mind experiments that examine whether children are able to reason about the beliefs of others, in particular when those beliefs differ from reality. We evaluate a number of recent neural models with memory augmentation. We find that all fail on our tasks, which require keeping track of inconsistent states of the world; moreover, the models' accuracy decreases notably when random sentences are introduced to the tasks at test.
Cite
CITATION STYLE
Nematzadeh, A., Burns, K., Grant, E., Gopnik, A., & Griffiths, T. L. (2018). Evaluating theory of mind in question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, EMNLP 2018 (pp. 2392–2400). Association for Computational Linguistics. https://doi.org/10.18653/v1/d18-1261
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.