Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can’t Answer?

5Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Question answering (QA)-giving correct answers to questions-is a popular task, but we test reverse question answering (RQA): for an input answer, give a question with that answer. Past work tests QA and RQA separately, but we test them jointly, comparing their difficulty, aiding benchmark design, and checking reasoning consistency. We run 16 LLMs on QA and RQA with trivia questions/answers, revealing: 1) Versus QA, LLMs are much less accurate in RQA for numerical answers, but slightly more accurate in RQA for textual answers; 2) LLMs often answer their own invalid questions from RQA accurately in QA, so RQA errors are not from knowledge gaps alone; 3) RQA errors correlate with question difficulty and inversely correlate with answer frequencies in the Dolma corpus; and 4) LLMs struggle to provide valid multi-hop questions. By finding question and answer types that lead to RQA errors, we suggest improvements for LLM reasoning.1

Cite

CITATION STYLE

APA

Balepur, N., Gu, F., Ravichander, A., Feng, S., Boyd-Graber, J., & Rudinger, R. (2025). Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can’t Answer? In Proceedings of the 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies: Long Papers, NAACL-HLT 2025 (Vol. 2, pp. 44–64). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.naacl-short.5

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free