Imputation Techniques in Machine Learning – A Survey

  • Angeline Christobel E
N/ACitations
Citations of this article
8Readers
Mendeley users who have this article in their library.

Abstract

Machine learning plays a pivotal role in data analysis and information extraction. However, one common challenge encountered in this process is dealing with missing values. Missing data can find its way into datasets for a variety of reasons. It can result from errors during data collection and management, intentional omissions, or even human errors. It's important to note that most machine learning models are not designed to handle missing values directly. Consequently, it becomes essential to perform data imputation before feeding the data into a machine learning model. Multiple techniques are available for imputing missing values, and the choice of technique should be made judiciously, considering various parameters. An inappropriate choice can disrupt the overall distribution of data values and subsequently impact the model's performance. In this paper, various imputation methods, including Mean, Median, K-nearest neighbors (KNN)-based imputation, Linear Regression, Miss Forest, and MICE are examined.

Cite

CITATION STYLE

APA

Angeline Christobel, E. al. (2023). Imputation Techniques in Machine Learning – A Survey. International Journal on Recent and Innovation Trends in Computing and Communication, 11(10), 1217–1221. https://doi.org/10.17762/ijritcc.v11i10.8662

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free