Abstract
We consider identification and estimation with an outcome missing not at random (MNAR). We study an identification strategy based on a so-called shadow variable . A shadow variable is assumed to be correlated with the outcome but independent of the missingness process conditional on the outcome and fully observed covariates. We describe a general condition for nonparametric identification of the full data law under MNAR using a valid shadow variable. Our condition is satisfied by many commonly used models; moreover, it is imposed on the complete cases, and therefore has testable implications with observed data only. We characterize the semiparametric efficiency bound for the class of regular and asymptotically linear estimators and derive a closed form for the efficient influence function. We describe a doubly robust and locally efficient estimation method and evaluate its performance on both simulation data and a real data example about home pricing. Problem statement Missingness not at random (MNAR) arises in many empirical studies in biomedical, socioeconomic, and epidemiological researches. A fundamental problem of MNAR is the identification problem, that is, the parameter of interest may not be uniquely determined with observed data. Besides, statistical inference is challenging under MNAR without identification. Methods This paper studies an identification strategy based on a so-called shadow variable. A shadow variable is assumed to be correlated with the outcome, but independent of the missingness process conditional on the outcome and fully observed covariates. A general condition for nonparametric identification of the full data law under MNAR using a valid shadow variable is provided. The corresponding semiparametric efficiency bound for the class of regular and asymptotically linear estimators is established. A doubly robust and locally efficient estimation method is proposed, evaluated on both simulation data, and applied to a real data example about home pricing. Results The proposed identification condition is satisfied by many commonly-used models; moreover, it is imposed on the complete cases, and therefore has testable implications with observed data only. The closed form for the efficient influence function is obtained, which motivates a doubly robust and locally efficient estimator. The estimator remains consistent even if certain working model is misspecified and attains the semiparametric efficiency bound if all working models are correct. Significance The paper describes the largest class of nonparametric models that are identifiable by the shadow variable approach, and establishes the semiparametric theory for this model. A novel doubly robust and locally efficient approach for the analysis of nonignorable missing data with a shadow variable is provided.
Cite
CITATION STYLE
Miao, W., Liu, L., Li, Y., Tchetgen Tchetgen, E. J., & Geng, Z. (2024). Identification and Semiparametric Efficiency Theory of Nonignorable Missing Data with a Shadow Variable. ACM / IMS Journal of Data Science, 1(2), 1–23. https://doi.org/10.1145/3592389
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.