Defining Data Science by a Data-Driven Quantification of the Community

33Citations
Citations of this article
30Readers
Mendeley users who have this article in their library.

Abstract

Data science is a new academic field that has received much attention in recent years. One reason for this is that our increasingly digitalized society generates more and more data in all areas of our lives and science and we are desperately seeking for solutions to deal with this problem. In this paper, we investigate the academic roots of data science. We are using data of scientists and their citations from Google Scholar, who have an interest in data science, to perform a quantitative analysis of the data science community. Furthermore, for decomposing the data science community into its major defining factors corresponding to the most important research fields, we introduce a statistical regression model that is fully automatic and robust with respect to a subsampling of the data. This statistical model allows us to define the ‘importance’ of a field as its predictive abilities. Overall, our method provides an objective answer to the question ‘What is data science?’.

Cite

CITATION STYLE

APA

Emmert-Streib, F., & Dehmer, M. (2019). Defining Data Science by a Data-Driven Quantification of the Community. Machine Learning and Knowledge Extraction, 1(1), 235–251. https://doi.org/10.3390/make1010015

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free