Utilization of synergetic human-machine clouds: A big data cleaning case

4Citations
Citations of this article
36Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Cloud computing and crowdsourcing are growing trends in IT. Combining the strengths of both machine and human clouds within a hybrid design enables us to overcome certain problems and achieve efficiencies. In this paper we present a case in which we developed a hybrid, throw-Away prototype software system to solve a big data cleaning problem in which we corrected and normalized a data set of 53,822 academic publication records. The first step in our solution consists of utilization of external DOI query web services to label the records with matching DOIs. Then we used customized string similarity calculation algorithms based on Levensthein Distance and Jaccard Index to grade the similarity between records. Finally we used crowdsourcing to identify duplicates among the residual record set consisting of similar yet not identical records. We consider this proof of concept to be successful and report that we achieved certain results that we could not have achieved by using either human or machine clouds alone.

Cite

CITATION STYLE

APA

Iren, D., Kul, G., & Bilgen, S. (2014). Utilization of synergetic human-machine clouds: A big data cleaning case. In 1st International Workshop on CrowdSourcing in Software Engineering, CSI-SE 2014 - Proceedings (pp. 15–18). Association for Computing Machinery, Inc. https://doi.org/10.1145/2593728.2593733

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free