Efficient Training Corpus Retrieval for Large Language Model Fine Tuning: A Case Study in Cancer

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The objective is to create an automated knowledge extraction tool for cancer research that builds high-quality academic corpora for LLM fine-tuning while investigating its effectiveness in interleukin-6 and bladder cancer domains. To address the current gap in knowledge retrieval techniques for cancer research data collection, we propose KnowledgePipeline, a novel automated tool that incorporates diverse aspects of academic papers and metadata. Our tool integrates content, co-citations, and co-authorship networks to construct domain-specific academic corpora suitable for fine-tuning LLMs. We leverage two LLMs (GPTJ-6.7B and Galactica30B) trained on domain-specific question-answer pairs from the refined data. The system's evaluation focuses on both the quality of extracted knowledge and the performance of fine-tuned models in open-ended question-answering tasks. We see that KnowledgePipeline offers a scalable, automated framework for domain-specific knowledge retrieval and fine-tuned applications in cancer research, advancing literature discovery and addressing critical biomedical challenges. It achieved high relevance scores of 68% for IL-6 and 74.5% for bladder cancer, with a fine-tuned Galactica-30B model demonstrating promising capabilities.

Cite

CITATION STYLE

APA

Das, A., Diala, C., Chen, G., Li, Z., Li, R., Anjum, O., & Zheng, W. J. (2025). Efficient Training Corpus Retrieval for Large Language Model Fine Tuning: A Case Study in Cancer. In Studies in Health Technology and Informatics (Vol. 329, pp. 1251–1255). IOS Press BV. https://doi.org/10.3233/SHTI251039

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free