Learning to crawl: Comparing classification schemes

132Citations
Citations of this article
88Readers
Mendeley users who have this article in their library.
Get full text

Abstract

Topical crawling is a young and creative area of research that holds the promise of benefiting from several sophisticated data mining techniques. The use of classification algorithms to guide topical crawlers has been sporadically suggested in the literature. No systematic study, however, has been done on their relative merits. Using the lessons learned from our previous crawler evaluation studies, we experiment with multiple versions of different classification schemes. The crawling process is modeled as a parallel best-first search over a graph defined by the Web. The classifiers provide heuristics to the crawler thus biasing it towards certain portions of the Web graph. Our results show that Naive Bayes is a weak choice for guiding a topical crawler when compared with Support Vector Machine or Neural Network. Further, the weak performance of Naive Bayes can be partly explained by extreme skewness of posterior probabilities generated by it. We also observe that despite similar performances, different topical crawlers cover subspaces on the Web with low overlap. © 2005 ACM.

Cite

CITATION STYLE

APA

Pant, G., & Srinivasan, P. (2005). Learning to crawl: Comparing classification schemes. ACM Transactions on Information Systems, 23(4), 430–462. https://doi.org/10.1145/1095872.1095875

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free