Empirical Evaluation of a Domain-Specific RAG-Based Conversational Assistant for Payroll Management: A Mixed-Methods Study of Accuracy, Satisfaction, and Usability

0Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

The integration of domain-specific conversational assistants into administrative and payroll information systems has the potential to improve operational efficiency and decision support in regulated enterprise environments. However, empirical evidence regarding their behavior, perceived usefulness, and operational performance in real-world administrative workflows remains limited. This paper introduces the design and empirical evaluation of a document-grounded conversational assistant for payroll management, implemented using a Retrieval-Augmented Generation (RAG) architecture enhanced with an evidence-grounded response validation layer designed to constrain responses to retrieved payroll documentation, enforce source traceability, and reduce unsupported content generation in administrative workflows. The primary evaluation dataset consisted of 200 controlled interactions generated by 20 experienced payroll administrators who completed predefined payroll-related tasks. In addition, an auxiliary dataset of approximately 800 supplementary interactions, produced during iterative expert testing and operational system validation activities, contributed to the broader operational interaction corpus used for exploratory robustness analysis and deployment-oriented system evaluation. The collected interaction logs were analyzed as an anonymized operational interaction corpus combining controlled participant sessions and supplementary expert-assisted testing activities. Consequently, the reported statistical findings should be interpreted primarily as an anonymized operational interaction dataset rather than strictly controlled inferential evidence. Quantitative assessment included response latency, user-perceived accuracy and satisfaction, and automatic evaluation metrics (BLEU, ROUGE-L, and semantic similarity), complemented by qualitative feedback collected through structured user sessions. The results indicate moderate levels of user-perceived accuracy and satisfaction. Semantic similarity demonstrates strong correlations with human evaluations of response quality, showing stronger associations than traditional n-gram metrics, while response latency shows only weak association with satisfaction within acceptable operational thresholds. The findings provide preliminary evidence that evidence-grounded validation mechanisms may support response traceability and reduction in unsupported content generation in administrative environments. This research contributes a domain-specific RAG-based conversational assistant incorporating evidence-grounded response validation mechanisms, together with an exploratory operational evaluation framework for assessing conversational AI behavior in compliance-sensitive enterprise environments.

Cite

CITATION STYLE

APA

Mitroulias, D., & Sioutas, S. (2026). Empirical Evaluation of a Domain-Specific RAG-Based Conversational Assistant for Payroll Management: A Mixed-Methods Study of Accuracy, Satisfaction, and Usability. Big Data and Cognitive Computing, 10(6). https://doi.org/10.3390/bdcc10060193

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free