User browsing behavior-driven web crawling

8Citations
Citations of this article
31Readers
Mendeley users who have this article in their library.
Get full text

Abstract

To optimize the performance of web crawlers, various page importance measures have been studied to select and order URLs in crawling. Most sophisticated measures (e.g. breadth-first and PageRank) are based on link structure. In this paper, we treat the problem from another perspective and propose to measure page importance through mining user interest and behaviors from web browse logs. Unlike most existing approaches which work on single URL, in this paper, both the log mining and the crawl ordering are performed at the granularity of URL pattern. The proposed URL pattern-based crawl orderings are capable to properly predict the importance of newly created (unseen) URLs. Promising experimental results proved the feasibility of our approach. © 2011 ACM.

Cite

CITATION STYLE

APA

Liu, M., Cai, R., Zhang, M., & Zhang, L. (2011). User browsing behavior-driven web crawling. In International Conference on Information and Knowledge Management, Proceedings (pp. 87–92). https://doi.org/10.1145/2063576.2063593

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free