The process of lexical blending is difficult to reliably predict. This difficulty has been shown by machine learning approaches in blend modeling, including attempts using then state-of-the-art LSTM deep neural networks trained on character embeddings, which were able to predict lexical blends given the ordered constituent words in less than half of cases, at maximum. This project introduces a novel model architecture which dramatically increases the correct prediction rates for lexical blends, using only Polynomial regression and Random Forest models. This is achieved by generating multiple possible blend candidates for each input word pairing and evaluating them based on observable linguistic features. The success of this model architecture illustrates the potential usefulness of observable linguistic features for problems that elude more advanced models which utilize only features discovered in the latent space.
CITATION STYLE
Saunders, J. (2023). Improving Automated Prediction of English Lexical Blends Through the Use of Observable Linguistic Features. In Proceedings of the Annual Meeting of the Association for Computational Linguistics (pp. 93–97). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2023.sigmorphon-1.10
Mendeley helps you to discover research relevant for your work.