Abstract
PLSA(Probabilistic Latent Semantic Analysis) is a popular topic modeling technique for exploring document collections. Due to the increasing prevalence of large datasets, there is a need to improve the scalability of computation in PLSA. In this paper, we propose a parallel PLSA algorithm called PPLSA to accommodate large corpus collections in the MapReduce framework. Our solution efficiently distributes computation and is relatively simple to implement. © 2012 IFIP International Federation for Information Processing.
Author supplied keywords
Cite
CITATION STYLE
Li, N., Zhuang, F., He, Q., & Shi, Z. (2012). PPLSA: Parallel probabilistic latent semantic analysis based on MapReduce. In IFIP Advances in Information and Communication Technology (Vol. 385 AICT, pp. 40–49). https://doi.org/10.1007/978-3-642-32891-6_8
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.