A performance of comparative study for semi-structured web data extraction model

3Citations
Citations of this article
5Readers
Mendeley users who have this article in their library.

Abstract

The extraction of information from multi-sources of web is an essential yet complicated step for data analysis in multiple domains. In this paper, we present a data extraction model based on visual segmentation, DOM tree and JSON approach which is known as Wrapper Extraction of Image using DOM and JSON (WEIDJ) for extracting semi-structured data from biodiversity web. The large number of information from multiple sources of web which is image’s information will be extracted using three different approach; Document Object Model (DOM), Wrapper image using Hybrid DOM and JSON (WHDJ) and Wrapper Extraction of Image using DOM and JSON (WEIDJ). Experiments were conducted on several biodiversity website. The experiment results show that WEIDJ approach promising results with respect to time analysis values. WEIDJ wrapper has successfully extracted greater than 100 images of data from the multi-source web biodiversity of over 15 different websites.

Cite

CITATION STYLE

APA

Sabri, I. A. A., & Man, M. (2019). A performance of comparative study for semi-structured web data extraction model. International Journal of Electrical and Computer Engineering, 9(6), 5463–5470. https://doi.org/10.11591/ijece.v9i6.pp5463-5470

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free