Probabilistic approaches to overcome content heterogeneity in data integration: A study case in systematic lupus erythematosus

1Citations
Citations of this article
10Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Integrating data from different sources into homogeneous dataset increases the opportunities to study human health. However, disparate data collections are often heterogeneous, which complicates their integration. In this paper, we focus on the issue of content heterogeneity in data integration. Traditional approaches for resolving content heterogeneity map all source datasets to a common data model that includes only shared data items, and thus omit all items that vary between datasets. Based on an example of three datasets in Systemic Lupus Erythematosus, we describe and experimentally evaluate a probabilistic data integration approach which propagates the uncertainty resulting from content heterogeneity into statistical inference, avoiding the need to map to a common data model.

Cite

CITATION STYLE

APA

Sampri, A., Geifman, N., Le Sueur, H., Doherty, P., Couch, P., Bruce, I., & Peek, N. (2020). Probabilistic approaches to overcome content heterogeneity in data integration: A study case in systematic lupus erythematosus. In Studies in Health Technology and Informatics (Vol. 270, pp. 387–391). IOS Press. https://doi.org/10.3233/SHTI200188

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free