Term Candidate Generation to Enrich Clinical Terminologies with Large Language Models

0Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Annotated language resources derived from clinical routine documentation form an intriguing asset for secondary use case scenarios. In this investigation, we report on how such a resource can be leveraged to identify additional term candidates for a chosen set of ICD-10 codes. We conducted a log-likelihood analysis, considering the co-occurrence of approximately 1.9 million de-identified ICD-10 codes alongside corresponding brief textual entries from problem lists in German. This analysis aimed to identify potential candidates with statistical significance set at p < 0.01, which were used as seed terms to harvest additional candidates by interfacing to a large language model in a second step. The proposed approach can identify additional term candidates at suitable performance values: hypernyms MAP@5=0.801, synonyms MAP@5 = 0.723 and hyponyms MAP@5 = 0.507. The re-use of existing annotated clinical datasets, in combination with large language models, presents an interesting strategy to bridge the lexical gap in standardized clinical terminologies and real-world jargon.

Cite

CITATION STYLE

APA

Kugic, A., Schulz, S., & Kreuzthaler, M. (2024). Term Candidate Generation to Enrich Clinical Terminologies with Large Language Models. In Studies in Health Technology and Informatics (Vol. 316, pp. 695–699). IOS Press BV. https://doi.org/10.3233/SHTI240509

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free