QARR-QGF: A Dual-Module Data Augmentation Framework for Enhanced Few-Shot Question Answering

0Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Few-shot question answering (QA) aims to train effective QA models using very limited annotated data, but existing approaches often struggle to generalize due to insufficient training examples. This paper proposes a two-stage data augmentation framework, called Question-Answer Replacement and Removal and Question Generation and Filtering (QARR-QGF), to address this challenge. The QARR module enhances pretraining data by systematically replacing and removing question-answer pairs to create more diverse examples. During fine-tuning, the QGF module generates paraphrased questions and applies semantic filtering to retain high-quality training samples. The framework is evaluated using three widely used generative models: Longformer-Encoder-Decoder (LED), BART, and T5, on the SQuAD, HotpotQA, and Natural Questions datasets. Experimental results show that QARR-QGF consistently improves performance across all datasets and few-shot settings. For example, the QARR-QGF-T5 model achieves F1 scores of 82.3% on SQuAD, 59.9% on HotpotQA, and 59.0% on Natural Questions in the 16-shot setting, outperforming previous state-of-the-art methods. These results demonstrate the effectiveness of QARR-QGF in improving few-shot QA performance by generating richer and more diverse training data.

Cite

CITATION STYLE

APA

Tan, S. W., Lee, C. P., Lim, K. M., & Alqahtani, A. (2025). QARR-QGF: A Dual-Module Data Augmentation Framework for Enhanced Few-Shot Question Answering. IEEE Access, 13, 160722–160736. https://doi.org/10.1109/ACCESS.2025.3609472

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free