Abstract
A reliable resume-job matching system helps a company recommend suitable candidates from a pool of resumes and helps a job seeker find relevant jobs from a list of job posts. However, since job seekers apply only to a few jobs, interaction labels in resume-job datasets are sparse. We introduce CONFIT V2, an improvement over CONFIT to tackle this sparsity problem. We propose two techniques to enhance the encoder's contrastive training process: augmenting job data with hypothetical reference resume generated by a large language model; and creating high-quality hard negatives from unlabeled resume/job pairs using a novel hard-negative mining strategy. We evaluate CONFIT V2 on two real-world datasets and demonstrate that it outperforms CONFIT and prior methods (including BM25 and OpenAI text-embedding-003), achieving an average absolute improvement of 13.8% in recall and 17.5% in nDCG across job-ranking and resume-ranking tasks.
Cite
CITATION STYLE
Yu, X., Xu, R., Xue, C., Zhang, J., Ma, X., & Yu, Z. (2025). CONFIT V2: Improving Resume-Job Matching using Hypothetical Resume Embedding and Runner-Up Hard-Negative Mining. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 12775–12790). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-acl.661
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.