Abstract
Due to the explosion in the size of the WWW1,4,5 it becomes essential to make the crawling process parallel. In this paper we present an architecture for a parallel crawler that consists of multiple crawling processes called as C-procs which can run on network of workstations. The proposed crawler is scalable, is resilient against system crashes and other event. The aim of this architecture is to efficiently and effectively crawl the current set of publically indexable web pages so that we can maximize the download rate while minimizing the overhead from parallelizatio
Cite
CITATION STYLE
Sharma, S., Sharma, A. K., & Gupta, J. P. (2011). A Novel Architecture of a Parallel Web Crawler. International Journal of Computer Applications, 14(4), 38–42. https://doi.org/10.5120/1846-2476
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.