Abstract
Hidden Markov models (HMMs) are valuable for their ability to provide exact and tractable inference. However, learning an HMM in an unsupervised manner involves a non-convex optimization problem that is plagued by poor local optima. Recent work on scaling HMMs has shown this challenge only intensifies as the number of hidden states grows. We provide a comprehensive empirical analysis of two approaches to enhance HMM optimization: reparameterization and initialization of HMM transition and emission parameters using neural networks. Through extensive experiments on language modeling, we find that (1) these techniques enable effective training of large-scale HMMs, (2) simple linear reparameterizations of HMM parameters perform as well as more complex neural ones, and (3) the two approaches are complementary, yielding the best results when combined.
Cite
CITATION STYLE
Lee, I., & Berg-Kirkpatrick, T. (2025). Optimizing Hidden Markov Language Models: An Empirical Study of Reparameterization and Initialization Techniques. In 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025 (pp. 7727–7738). Association for Computational Linguistics (ACL). https://doi.org/10.18653/v1/2025.findings-naacl.429
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.