Misspelling Detection from Noisy Product Images

0Citations
Citations of this article
67Readers
Mendeley users who have this article in their library.

Abstract

Misspellings are introduced on products either due to negligence or as an attempt to deliberately deceive stakeholders. This leads to a revenue loss for online sellers and fosters customer mistrust. Existing spelling research has primarily focused on advancement in misspelling correction and the approach for misspelling detection has remained the use of a large dictionary. The dictionary lookup results in the incorrect detection of several non-dictionary words as misspellings. In this paper, we propose a method to automatically detect misspellings from product images in an attempt to reduce false positive detections. We curate a large scale corpus, define a rich set of features and propose a novel model that leverages importance weighting to account for within class distributional variance. Finally, we experimentally validate this approach on both the curated corpus and an out-of-domain public dataset and show that it leads to a relative improvement of up to 20% in F1 score. The approach thus creates a more robust, generalized deployable solution and reduces reliance on large scale custom dictionaries used today.

Cite

CITATION STYLE

APA

Rao, V. N., & Shen, M. (2020). Misspelling Detection from Noisy Product Images. In COLING 2020 - 28th International Conference on Computational Linguistics, Proceedings of the Industry Track (pp. 124–135). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2020.coling-industry.12

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free