Integrating Clustering and Multi-Document Summarization by Bi-Mixture Probabilistic Latent Semantic Analysis (PLSA) with Sentence Bases

Chao Shen; Tao Li; Chris H.Q. Ding

Conference ProceedingsOPEN ACCESS

Integrating Clustering and Multi-Document Summarization by Bi-Mixture Probabilistic Latent Semantic Analysis (PLSA) with Sentence Bases

Proceedings of the 25th AAAI Conference on Artificial Intelligence, AAAI 2011 (2011) 914-920

DOI: 10.1609/aaai.v25i1.7977

14Citations

22Readers

Abstract

Probabilistic Latent Semantic Analysis (PLSA) has been popularly used in document analysis. However, as it is currently formulated, PLSA strictly requires the number of word latent classes to be equal to the number of document latent classes. In this paper, we propose Bi-mixture PLSA, a new formulation of PLSA that allows the number of latent word classes to be different from the number of latent document classes. We further extend Bi-mixture PLSA to incorporate the sentence information, and propose Bi-mixture PLSA with sentence bases (Bi-PLSAS) to simultaneously cluster and summarize the documents utilizing the mutual influence of the document clustering and summarization procedures. Experiments on real-world datasets demonstrate the effectiveness of our proposed methods.

Cite

CITATION STYLE

APA

Shen, C., Li, T., & Ding, C. H. Q. (2011). Integrating Clustering and Multi-Document Summarization by Bi-Mixture Probabilistic Latent Semantic Analysis (PLSA) with Sentence Bases. In Proceedings of the 25th AAAI Conference on Artificial Intelligence, AAAI 2011 (pp. 914–920). AAAI Press. https://doi.org/10.1609/aaai.v25i1.7977

Integrating Clustering and Multi-Document Summarization by Bi-Mixture Probabilistic Latent Semantic Analysis (PLSA) with Sentence Bases

Abstract

Cite

Register to see more suggestions