Parallel crawling for online social networks

Duen Horng Chau; Shashank Pandit; Samuel Wang; Christos Faloutsos

Conference Proceedings

Parallel crawling for online social networks

16th International World Wide Web Conference, WWW2007 (2007) 1283-1284

DOI: 10.1145/1242572.1242809

84Citations

90Readers

Get full text

Abstract

Given a huge online social network, how do we retrieve information from it through crawling? Even better, how do we improve the crawling performance by using parallel crawlers that work independently? In this paper, we present the framework of parallel crawlers for online social networks, utilizing a centralized queue. To show how this works in practice, we describe our implementation of the crawlers for an online auction website. The crawlers work independently, therefore the failing of one crawler does not affect the others at all. The framework ensures that no redundant crawling would occur. Using the crawlers that we built, we visited a total of approximately 11 million auction users, about 66,000 of which were completely crawled.

Author supplied keywords

Cite

CITATION STYLE

APA

Chau, D. H., Pandit, S., Wang, S., & Faloutsos, C. (2007). Parallel crawling for online social networks. In 16th International World Wide Web Conference, WWW2007 (pp. 1283–1284). https://doi.org/10.1145/1242572.1242809

Parallel crawling for online social networks

Abstract

Author supplied keywords

Cite

Register to see more suggestions