Data Cleaning in Text File

  • Bhattacharjee A
N/ACitations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

Data cleaning is an automated process of detecting, removing and correcting incomplete, incorrect, inaccurate and irrelevant data from a record set. Our system works on simple text (*.txt) files using Extract, Transform and Load (ETL) model. In this paper we present a set of algorithms to correct errors such as alpha- numeric errors, invalid gender, invalid ID pattern and redundant ID error. The text files are used as data storage which stores data in a tabular format and the algorithms are applied on each field value depending on its nature.

Cite

CITATION STYLE

APA

Bhattacharjee, A. K. (2013). Data Cleaning in Text File. IOSR Journal of Computer Engineering, 9(2), 17–21. https://doi.org/10.9790/0661-0921721

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free