Effective processing of unstructured data using python in Hadoop map reduce

  • Kousalya K
  • Javed Parvez S
N/ACitations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

In present scenario, the growing data are naturally unstructured. In this case to handle the wide range of data, is difficult. The proposed paper is to process the unstructured text data effectively in Hadoop map reduce using Python. Apache Hadoop is an open source platform and it widely uses Map Reduce framework. Map Reduce is popular and effective for processing the unstructured data in parallel manner.  There are two stages in map reduce, namely transform and repository. Here the input splits into small blocks and worker node process individual blocks in parallel. This map reduce generally is based on java. While Hadoop Streaming allows writing mapper and reducer in other languages like Python. In this paper, we are going to show an alternative way of processing the growing unstructured content data by using python. We will also compare the performance between java based and non-java based programs.

Cite

CITATION STYLE

APA

Kousalya, K., & Javed Parvez, S. (2018). Effective processing of unstructured data using python in Hadoop map reduce. International Journal of Engineering & Technology, 7(2.21), 417. https://doi.org/10.14419/ijet.v7i2.21.12456

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free