Abstract
Few-shot question answering (QA) aims to train effective QA models using very limited annotated data, but existing approaches often struggle to generalize due to insufficient training examples. This paper proposes a two-stage data augmentation framework, called Question-Answer Replacement and Removal and Question Generation and Filtering (QARR-QGF), to address this challenge. The QARR module enhances pretraining data by systematically replacing and removing question-answer pairs to create more diverse examples. During fine-tuning, the QGF module generates paraphrased questions and applies semantic filtering to retain high-quality training samples. The framework is evaluated using three widely used generative models: Longformer-Encoder-Decoder (LED), BART, and T5, on the SQuAD, HotpotQA, and Natural Questions datasets. Experimental results show that QARR-QGF consistently improves performance across all datasets and few-shot settings. For example, the QARR-QGF-T5 model achieves F1 scores of 82.3% on SQuAD, 59.9% on HotpotQA, and 59.0% on Natural Questions in the 16-shot setting, outperforming previous state-of-the-art methods. These results demonstrate the effectiveness of QARR-QGF in improving few-shot QA performance by generating richer and more diverse training data.
Author supplied keywords
Cite
CITATION STYLE
Tan, S. W., Lee, C. P., Lim, K. M., & Alqahtani, A. (2025). QARR-QGF: A Dual-Module Data Augmentation Framework for Enhanced Few-Shot Question Answering. IEEE Access, 13, 160722–160736. https://doi.org/10.1109/ACCESS.2025.3609472
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.