TrialBench: Multi-Modal AI-Ready Datasets for Clinical Trial Prediction

N/ACitations
Citations of this article
17Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Clinical trials are pivotal for developing new medical treatments but typically carry risks such as patient mortality and enrollment failure that waste immense efforts spanning over a decade. Applying artificial intelligence (AI) to predict key events in clinical trials holds great potential for providing insights to guide trial designs. However, complex data collection and question definition requiring medical expertise have hindered the involvement of AI thus far. This paper tackles these challenges by presenting a comprehensive suite of 23 meticulously curated AI-ready datasets covering multi-modal input features and 8 crucial prediction challenges in clinical trial design, encompassing prediction of trial duration, patient dropout rate/event, serious adverse event, mortality event, trial approval outcome, trial failure reason, drug dose, and design of eligibility criteria. Furthermore, we provide basic validation methods for each task to ensure the datasets’ usability and reliability. We anticipate that the availability of such open-access datasets will catalyze the development of advanced AI approaches for clinical trial design, ultimately advancing clinical trial research and accelerating medical solution development.

Cite

CITATION STYLE

APA

Chen, J., Hu, Y., Cai, M., Lu, Y., Wang, Y., Cao, X., … Fu, T. (2025). TrialBench: Multi-Modal AI-Ready Datasets for Clinical Trial Prediction. Scientific Data , 12(1). https://doi.org/10.1038/s41597-025-05680-8

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free