CopyCat: Near-Duplicates within and between the ClueWeb and the Common Crawl

20Citations
Citations of this article
11Readers
Mendeley users who have this article in their library.
Get full text

Abstract

The amount of near-duplicates in web crawls like the ClueWeb or Common Crawl demands from their users either to develop a preprocessing pipeline for deduplication, which is costly both computationally and in person hours, or accepting the undesired effects that near-duplicates have on reliability and validity of experiments. We introduce ChatNoir-CopyCat-21, which simplifies deduplication significantly. It comes in two parts: (1) A compilation of near-duplicate documents within the ClueWeb09, the ClueWeb12, and two Common Crawl snapshots, as well as between selections of these crawls, and (2) a software library that implements the deduplication of arbitrary document sets. Our analysis shows that 14 - 52, of the documents within a crawl and around0.7 - 2.5, between the crawls are near-duplicates. Two showcases demonstrate the application and usefulness of our resource.

Cite

CITATION STYLE

APA

Fröbe, M., Bevendorff, J., Gienapp, L., Völske, M., Stein, B., Potthast, M., & Hagen, M. (2021). CopyCat: Near-Duplicates within and between the ClueWeb and the Common Crawl. In SIGIR 2021 - Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 2398–2404). Association for Computing Machinery, Inc. https://doi.org/10.1145/3404835.3463246

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free