Candidate Profile Summarization- A RAG Approach with Synthetic Data Generation for Tech Jobs

0Citations
Citations of this article
7Readers
Mendeley users who have this article in their library.

Abstract

As Large Language Models (LLMs) become increasingly applied to resume evaluation and candidate selection, this study investigates the effectiveness of using in-context example resumes to generate synthetic data. We compare a Retrieval-Augmented Generation (RAG) system to a Named Entity Recognition (NER)based baseline for job-resume matching, generating diverse synthetic resumes with models like Mixtral-8x22B-Instruct-v0.1. Our results show that combining BERT, ROUGE, and Jaccard similarity metrics effectively assesses synthetic resume quality, ensuring the least lexical overlap along with high similarity and diversity. Our experiments show that RAG notably outperforms NER for retrieval tasks—though generation-based summarization remains challenged by role differentiation. Human evaluation further highlights issues of factual accuracy and completeness, emphasizing the importance of in-context examples, prompt engineering, and improvements in summary generation for robust, automated candidate selection.

Cite

CITATION STYLE

APA

Afzal, A., Subedi, I., & Matthes, F. (2025). Candidate Profile Summarization- A RAG Approach with Synthetic Data Generation for Tech Jobs. In International Conference Recent Advances in Natural Language Processing, RANLP (pp. 22–31). Incoma Ltd. https://doi.org/10.26615/978-954-452-098-4-003

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free