Assessing data extraction in randomized clinical trials with large language models

1Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Background: Data extraction is an essential step in evidence synthesis but remains time-consuming and prone to human error. Large language models (LLMs) such as ChatGPT-4 and Claude 3 Opus may offer partial automation solutions. This proof-of-concept study evaluated their preliminary performance in extracting data from full-text randomized controlled trial (RCT) reports within systematic reviews. Methods: Two previously validated systematic reviews published in European Urology (105 trials in total) were selected. Standardized prompts were created and optimized with ChatGPT-4 and tested independently across trials using both ChatGPT-4 (paid version) and Claude 3 Opus via the standard web interface. Each prompt was executed three times, and only the first output was used to calculate the proportion of correctly extracted items (Pacc). Extracted data were compared with verified gold-standard datasets. Results: For binary outcomes, ChatGPT-4 and Claude 3 Opus showed high accuracy in group size extraction (Pacc: 91%-94%) and moderate accuracy for event counts (Pacc: 57%-71%). For continuous outcomes, group size accuracy was moderate (Pacc: 59%-69%), while mean and standard deviation extraction was poor (Pacc: 24%-56%). Test-retest reliability was substantial to almost perfect. Conclusions: Current LLMs can assist in automating data extraction for binary outcomes but remain inconsistent for continuous outcomes. These preliminary results should be interpreted with caution. Further research using larger datasets and iterative prompt refinement is needed before LLMs can be reliably integrated into systematic review workflows.

Cite

CITATION STYLE

APA

Yisha, Z., Zou, P., Li, S., Zhang, L., Guo, L., Gu, A., … Wang, X. (2026). Assessing data extraction in randomized clinical trials with large language models. BMC Medical Research Methodology, 26(1). https://doi.org/10.1186/s12874-025-02729-5

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free