Building scalable web archives

1Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

This paper aims at introducing the Internet Memory Foundation platform based on its distributed infrastructure and the associated tools and workflows that facilitate data management and preservation actions at large scale. IMF's main concern over the past years has been related to scalability issues in terms of crawling, indexing, preserving and accessing content. To answer these issues, the Foundation developed its own crawler and built a new infrastructure. This paper aims at presenting our infrastructure and crawler and at sharing challenges met while building them as well as the approach taken to solve preservation issues inherent to scalable archives. It will also highlight new horizons arising for web archives in relation to analytics use cases. © 2014 Society for Imaging Science and Technology.

Cite

CITATION STYLE

APA

Medjkoune, L., Barton, S., Carpentier, F., Masanès, J., & Pop, R. (2014). Building scalable web archives. In Archiving 2014 - Final Program and Proceedings (pp. 138–143). Society for Imaging Science and Technology. https://doi.org/10.2352/issn.2168-3204.2014.11.1.art00030

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free