Abstract
Similarity-based language models can be used to deal with data sparseness. A general method for using similarity-based models is proposed to improve the estimates of existing language models. In a pilot study, the language modeling performance of a similarity-based model with a standard back-off model is compared. Following this, a more detailed study is performed where several similarity-based models and parameter settings are compared on a smaller, more manageable word-sense disambiguation task. The main observation is that the similarity-based methods perform much better on unseen word pairs.
Cite
CITATION STYLE
Dagan, I., Lee, L., & Pereira, F. G. N. (1999). Similarity-based models of word cooccurrence probabilities. Machine Learning, 34(1), 43–69. https://doi.org/10.1023/a:1007537716579
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.