Abstract
This paper investigates the problem of text normalisation; specifically, the normalisation of non-standard words (NSWs) in English. Non-standard words can be defined as those word tokens which do not have a dictionary entry, and cannot be pronounced using the usual letterto- phoneme conversion rules; e.g. lbs, 99.3%, #EMNLP2017. NSWs pose a challenge to the proper functioning of textto- speech technology, and the solution is to spell them out in such a way that they can be pronounced appropriately. We describe our four-stage normalisation system made up of components for detection, classification, division and expansion of NSWs. Performance is favourabe compared to previous work in the field (Sproat et al. 2001, Normalization of nonstandard words), as well as state-of-the-art text-to-speech software. Further, we update Sproat et al.'s NSW taxonomy, and create a more customisable system where users are able to input their own abbreviations and specify into which variety of English (currently available: British or American) they wish to normalise.
Cite
CITATION STYLE
Flint, E., Ford, E., Thomas, O., Caines, A., & Buttery, P. (2017). A Text Normalisation System for Non-Standard English Words. In 3rd Workshop on Noisy User-Generated Text, W-NUT 2017 - Proceedings of the Workshop (pp. 107–115). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/w17-4414
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.