The NoisyOffice Database: A Corpus to Train Supervised Machine Learning Filters for Image Processing

10Citations
Citations of this article
9Readers
Mendeley users who have this article in their library.

Abstract

This paper presents the 'NoisyOffice' database. It consists of images of printed text documents with noise mainly caused by uncleanliness from a generic office, such as coffee stains and footprints on documents or folded and wrinkled sheets with degraded printed text. This corpus is intended to train and evaluate supervised learning methods for cleaning, binarization and enhancement of noisy images of grayscale text documents. As an example, several experiments of image enhancement and binarization are presented by using deep learning techniques. Also, double-resolution images are also provided for testing super-resolution methods. The corpus is freely available at UCI Machine Learning Repository. Finally, a challenge organized by Kaggle Inc. to denoise images, using the database, is described in order to show its suitability for benchmarking of image processing systems.

Cite

CITATION STYLE

APA

Castro-Bleda, M. J., España-Boquera, S., Pastor-Pellicer, J., & Zamora-Martínez, F. (2020). The NoisyOffice Database: A Corpus to Train Supervised Machine Learning Filters for Image Processing. Computer Journal, 63(11), 1658–1667. https://doi.org/10.1093/comjnl/bxz098

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free