Abstract
Automated resume evaluation for assessing candidate-job alignment remains challenging due to the complexity of matching qualifications across education, skills, and experience dimensions. While recent advances in large language models (LLMs) show promise, their effectiveness for domain-specific resume screening remains underexplored. We present a comprehensive benchmark comparing traditional machine learning, deep learning, pre-trained transformers, fine-tuned LLMs, and multi-agent architectures on professionally annotated multi-domain resumes spanning six job categories. Our proposed heterogeneous multi-agent framework—evaluating Education, Skills, and Experience via independently fine-tuned specialist models—achieves the best overall performance (MAE: 0.065, Pearson r: 0.896, R2: 0.766; best configuration: QwenFT+ QwenFT+ GemmaFT), outperforming both fine-tuned single-agent models (MAE: 0.064, R2: 0.729) and all five commercial state-of-the-art baselines in zero-shot settings: GPT-5 (MAE: 0.252), Gemini 3 (MAE: 0.128), Grok (Auto, xAI) (MAE: 0.121), Mistral Large (MAE: 0.1226), and Claude Sonnet 4.5 (MAE: 0.097). These results show that strategic assignment of domain-specifically fine-tuned models to evaluation sub-dimensions yields interpretable, component-level assessments while achieving superior accuracy over both monolithic and commercial alternatives.
Author supplied keywords
Cite
CITATION STYLE
Sagor Chowdhury, M., Chowdhury, A. F., Banu, A., Hossain, R., & Chowdhury, M. (2026). Automated Resume Evaluation Using Large Language Models: A Multi-Agent Framework With Fine-Tuning and Prompt Engineering. IEEE Access, 14, 84085–84102. https://doi.org/10.1109/ACCESS.2026.3696456
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.