On the use of gradient boosting methods to improve the estimation with data obtained with self-selection procedures

Luis Castro-Martín; María Del Mar Rueda; Ramón Ferri-García; César Hernando-Tamayo

Journal ArticleOPEN ACCESS

On the use of gradient boosting methods to improve the estimation with data obtained with self-selection procedures

Mathematics (2021) 9(23)

DOI: 10.3390/math9232991

11Citations

5Readers

Abstract

In the last years, web surveys have established themselves as one of the main methods in empirical research. However, the effect of coverage and selection bias in such surveys has undercut their utility for statistical inference in finite populations. To compensate for these biases, researchers have employed a variety of statistical techniques to adjust nonprobability samples so that they more closely match the population. In this study, we test the potential of the XGBoost algorithm in the most important methods for estimation that integrate data from a probability survey and a nonprobability survey. At the same time, a comparison is made of the effectiveness of these methods for the elimination of biases. The results show that the four proposed estimators based on gradient boosting frameworks can improve survey representativity with respect to other classic prediction methods. The proposed methodology is also used to analyze a real nonprobability survey sample on the social effects of COVID-19.

Author supplied keywords

Cite

CITATION STYLE

APA

Castro-Martín, L., Rueda, M. D. M., Ferri-García, R., & Hernando-Tamayo, C. (2021). On the use of gradient boosting methods to improve the estimation with data obtained with self-selection procedures. Mathematics, 9(23). https://doi.org/10.3390/math9232991

On the use of gradient boosting methods to improve the estimation with data obtained with self-selection procedures

Abstract

Author supplied keywords

Cite

Register to see more suggestions