Data augmentation based on large language models for radiological report classification

7Citations
Citations of this article
14Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The International Classification of Diseases (ICD) is fundamental in the field of healthcare as it provides a standardized framework for the classification and coding of medical diagnoses and procedures, enabling the understanding of international public health patterns and trends. However, manually classifying medical reports according to this standard is a slow, tedious and error-prone process, which shows the need for automated systems to offload the healthcare professional of this task and to reduce the number of errors. In this paper, we propose an automated classification system based on Natural Language Processing to analyze radiological reports and classify them according to the ICD-10. Since the specialized use of the language of radiological reports and the usual unbalanced distribution of medical report sets, we propose a methodology grounded in leveraging large language models for augmenting the data of unrepresented classes and adapting the classification language models to the specific use of the language of radiological reports. The results show that the proposed methodology enhances the classification performance on the CARES corpus of radiological reports.

Cite

CITATION STYLE

APA

Collado-Montañez, J., Martín-Valdivia, M. T., & Martínez-Cámara, E. (2025). Data augmentation based on large language models for radiological report classification. Knowledge-Based Systems, 308. https://doi.org/10.1016/j.knosys.2024.112745

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free