WARCProcessor: An integrative tool for building and management of web spam corpora

1Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

Abstract

In this work we present the design and implementation of WARCProcessor, a novel multiplatform integrative tool aimed to build scientific datasets to facilitate experimentation in web spam research. The developed application allows the user to specify multiple criteria that change the way in which new corpora are generated whilst reducing the number of repetitive and error prone tasks related with existing corpus maintenance. For this goal,WARCProcessor supports up to six commonly used data sources for web spam research, being able to store output corpus in standardWARC format together with complementary metadata files. Additionally, the application facilitates the automatic and concurrent download of web sites from Internet, giving the possibility of configuring the deep of the links to be followed as well as the behaviour when redirected URLs appear. WARCProcessor supports both an interactive GUI interface and a command line utility for being executed in background.

Cite

CITATION STYLE

APA

Callón, M., Fdez-Glez, J., Ruano-Ordás, D., Laza, R., Pavón, R., Fdez-Riverola, F., & Méndez, J. R. (2018). WARCProcessor: An integrative tool for building and management of web spam corpora. Sensors (Switzerland), 18(1). https://doi.org/10.3390/s18010016

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free