Abstract
The terms reproducibility and replicability have been used interchangeably by some scientific communities and by media, and with the opposite meanings by others, causing much confusion. The 2019 report on “Reproducibility and Replicability in Science” issued by US National Academies of Sciences, Engineering and Medicine (NASEM) made an important contribution to delineate the two terms by equating reproducibility with computational reproducibility and replicability with scientific replicability. However, neither of them in itself can guarantee reliability. Reliability does not imply absolute truth, but it does require that our findings can be triangulated, can pass reasonable stress tests and fair-minded sensitivity tests, and they do not contradict the best available theory and scientific understanding, unless the findings are designed to challenge the existing common wisdom. The quality of data and information plays far important roles than their quantity in ensuring reliability. This talk reflects on these issues based on my statistical research on quantifying quality of big data, and as the founding Editor-in-Chief of Harvard Data Science Review (HDSR), an experience that has provided me a much broader data science perspective. Along the way, using US election prediction and COVID-19 testing as two recent examples, I will demonstrate how small our big data are when we take into account their quality.
Cite
CITATION STYLE
Meng, X.-L. (2020). Reproducibility, Replicability, and Reliability. Harvard Data Science Review, 2(4). https://doi.org/10.1162/99608f92.dbfce7f9
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.