Abstract
In a robust information retrieval system, the documents should be represented by considering the variations of word distributions in different paragraphs or segments. A nonstationary latent Dirichlet allocation (NLDA) was established by incorporating a Markov chain to detect the stylistic segments in a heterogeneous document. Each segment corresponds to a particular style and is generated by different word distributions. However, such NLDA is constrained by a fixed number of segments for different lengths of documents. This paper presents a new adaptive segment model (ASM) by adaptively building the topic-based document model with different segment numbers. By incorporating a multinomial hidden variable with Dirichlet prior, the inference procedure of ASM parameters is built through a variational Bayes EM algorithm. In the experiments, the proposed ASM is evaluated for spoken document retrieval using TDT2 corpus. ASM achieves better performance than LDA and NLDA. ©2010 IEEE.
Author supplied keywords
Cite
CITATION STYLE
Chueh, C. H., & Chien, J. T. (2010). Adaptive segment model for spoken document retrieval. In 2010 7th International Symposium on Chinese Spoken Language Processing, ISCSLP 2010 - Proceedings (pp. 261–264). https://doi.org/10.1109/ISCSLP.2010.5684896
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.