Mining and representing unstructured nicotine use data in a structured format for secondary use

2Citations
Citations of this article
13Readers
Mendeley users who have this article in their library.

Abstract

The objective of this study was to use rules, NLP and machine learning for addressing the problem of clinical data interoperability across healthcare providers. Addressing this problem has the potential to make clinical data comparable, retrievable and exchangeable between healthcare providers. Our focus was in giving structure to unstructured patient smoking information. We collected our data from the MIMIC-III database. We wrote rules for annotating the data, then trained a CRF sequence classifier. We obtained an f-measure of 86%, 72%, 69%, 80%, and 12% for substance smoked, frequency, amount, temporal, and duration respectively. Amount smoked yielded a small value due to scarcity of related data. Then for smoking status we obtained an f-measure of 94.8% for non-smoker class, 83.0% for current-smoker, and 65.7% for past-smoker. We created a FHIR profile for mapping the extracted data based on openEHR reference models, however in future we will explore mapping to CIMI models.

Cite

CITATION STYLE

APA

Ngwenya, M., & Bankole, F. (2019). Mining and representing unstructured nicotine use data in a structured format for secondary use. In Proceedings of the Annual Hawaii International Conference on System Sciences (Vol. 2019-January, pp. 3751–3760). IEEE Computer Society. https://doi.org/10.24251/hicss.2019.453

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free