Enhancing Transparency and Replicability in Data Collection: Lessons from the Construction of Three Education Datasets

0Citations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Assembling datasets is crucial for advancing social science research, but researchers who construct datasets often face difficult decisions with little guidance. Once public, these datasets are sometimes used without proper consideration of their creators’ choices and how these affect the validity of inferences. To support both data creators and data users, we discuss the strengths, limitations, and implications of various data collection methodologies and strategies, showing how seemingly trivial methodological differences can significantly impact conclusions. The lessons we distill build on the process of constructing three cross-national datasets on education systems. Despite their common focus, these datasets differ in the dimensions they measure, as well as their definitions of key concepts, coding thresholds and other assumptions, types of coders, and sources. From these lessons, we develop and propose general guidelines for dataset creators and users aimed at enhancing transparency, replicability, and valid inferences in the social sciences.

Cite

CITATION STYLE

APA

Del Río, A., Kim, W., Knutsen, C. H., Neundorf, A., Paglayan, A. S., & Nazrullaeva, E. (2025). Enhancing Transparency and Replicability in Data Collection: Lessons from the Construction of Three Education Datasets. Perspectives on Politics. https://doi.org/10.1017/S1537592725103058

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free