Exploratory Data Mining for Subgroup Cohort Discoveries and Prioritization

25Citations
Citations of this article
61Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Finding small homogeneous subgroup cohorts in large heterogeneous populations is a critical process for hypothesis development in biomedical research. Concurrent computational approaches are still lacking in robust answers to the question 'what hypotheses are likely to be novel and to produce clinically relevant results with well thought-out study designs?' We have developed a novel subgroup discovery method which employs a deep exploratory mining process to slice and dice thousands of potential subpopulations and prioritize potential cohorts based on their explainable contrast patterns and which may provide interventionable insights. We conducted computational experiments on both synthesized data and a clinical autism data set to assess performance quantitatively for coverage of pre-defined cohorts and qualitatively for novel knowledge discovery, respectively. We also conducted a scaling analysis using a distributed computing environment to suggest computational resource needs for when the subpopulation number increases. This work will provide a robust data-driven framework to automatically tailor potential interventions for precision health.

Cite

CITATION STYLE

APA

Liu, D., Baskett, W., Beversdorf, D., & Shyu, C. R. (2020). Exploratory Data Mining for Subgroup Cohort Discoveries and Prioritization. IEEE Journal of Biomedical and Health Informatics, 24(5), 1456–1468. https://doi.org/10.1109/JBHI.2019.2939149

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free