Abstract
Retrieval-Augmented Generation (RAG) systems are increasingly adopted in government services, yet different administrations have varying customization needs and lack standardized methods to evaluate performance. In particular, general-purpose evaluation approaches fail to show how well a system meets domain-specific expectations. This paper presents CuBE (Customizable Bounds Evaluation), a tailored evaluation framework for RAG systems in public administration. CuBE integrates large language model (LLM) scoring, customizable evaluation dimensions, and a bounded scoring paradigm with baseline and upper-bound reference sets, enhancing fairness, consistency, and interpretability. We further introduce Lightweight Targeted Assessment (LTA) to support efficient customization. CuBE is validated on GSIA (Guizhou Provincial Government Service Center Intelligent Assistant) by using four state-of-the-art language models. The results show that CuBE produces robust, stable, and model-agnostic evaluations while reducing reliance on manual annotation and facilitating system optimization and rapid iteration. Moreover, CuBE informs parameter settings, enabling developers to design RAG systems that better meet customizer needs. This study establishes a replicable paradigm for trustworthy and efficient evaluation of RAG systems in complex government service scenarios.
Author supplied keywords
Cite
CITATION STYLE
Yang, B., Yu, X., Zheng, X., Nong, J., Liu, Z., Dai, X., & Xie, X. (2025). CuBE: A Customizable Bounds Evaluation Framework for Automated Assessment of RAG Systems in Government Services. Applied Sciences (Switzerland), 15(19). https://doi.org/10.3390/app151910447
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.