What, Why, and How: An Empiricist’s Guide to Double/Debiased Machine Learning

  • Shi B
  • Mao X
  • Yang M
  • et al.
N/ACitations
Citations of this article
12Readers
Mendeley users who have this article in their library.
Get full text

Abstract

We provide an introduction to double/debiased machine learning (DML), a framework that enables effect estimation when dealing with complex, high-dimensional data. In many empirical analyses, especially in fields such as information systems, researchers face difficult choices about which control variables to include and how to model their relationships with the outcome. These modeling decisions can significantly change results, leading to uncertainty about which findings are reliable. DML offers a practical solution by combining modern machine learning with rigorous statistical inference. The idea is to let flexible ML models (such as random forests or gradient boosting) capture complex relationships among control variables while still delivering reliable estimates for the key effect of interest. DML can be applied to many familiar research designs, including standard regression with controls, instrumental variables, difference in differences, and models that incorporate ML-generated features. Empirical studies and simulations show that DML is typically more robust to misspecification than traditional regression and more reliable than earlier semiparametric methods. However, DML is not automatic—it still requires sound research design and high-quality machine learning estimation. Used thoughtfully, DML provides a powerful, flexible, and statistically grounded approach for empirical research in modern data environments.This research commentary introduces double/debiased machine learning (DML), a novel methodological framework, to the information systems (IS) research community, demonstrating its power to address the challenges of empirical model specifications. DML combines the flexibility of modern machine learning (ML) techniques with the rigor of semiparametric statistical theory, enabling effective modeling of complex functions alongside valid statistical inference. The paper provides an accessible and comprehensive overview of DML’s key elements—Neyman orthogonality, cross-fitting, and high-quality ML estimation—and their roles in achieving methodological flexibility and rigor. The versatility of DML is illustrated through applications in several empirical settings common in IS research, including standard linear regression with control covariates, instrumental variable regressions, difference in differences, and scenarios with ML-generated covariates. Comparative simulations and real data analyses show that DML outperforms traditional parametric and semiparametric methods, and they also illustrate the importance of DML’s key elements. Finally, we highlight potential misconceptions and pitfalls in applying DML and offer practical advice for empirical researchers. Given the increasing complexity of data and research questions in the IS field, DML offers a timely and powerful tool for empirical researchers. By promoting a deeper understanding and appropriate use of DML, this commentary aims to empower empirical research in IS.History: Karthik Kannan, Senior Editor; Anuj Kumar, Associate Editor.Funding: X. Mao is supported in part by the National Natural Science Foundation of China [Grants 72201150, 72322001, and 72293561] and the National Key R&D Program of China [Grant 2022ZD0116700]. B. Li’s research was supported by the National Natural Science Foundation of China [Grants 72171131 and 72133002].Supplemental Material: The online appendix is available at https://doi.org/10.1287/isre.2024.0888 .

Cite

CITATION STYLE

APA

Shi, B., Mao, X., Yang, M., & Li, B. (2026). What, Why, and How: An Empiricist’s Guide to Double/Debiased Machine Learning. Information Systems Research, 37(2), 1259–1275. https://doi.org/10.1287/isre.2024.0888

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free