Abstract
Background: Establishing the relationship between microbiota and specific diseases is important but requires appropriate statistical methodology. A specialized feature of microbiome count data is the presence of a large number of zeros, which makes it difficult to analyze in case-control studies. Most existing approaches either add a small number called a pseudo-count or use probability models such as the multinomial and Dirichlet-multinomial distributions to explain the excess zero counts, which may produce unnecessary biases and impose a correlation structure taht is unsuitable for microbiome data. Results: The purpose of this article is to develop a new probabilistic model, called BERnoulli and MUltinomial Distribution-based latent Allocation (BERMUDA), to address these problems. BERMUDA enables us to describe the differences in bacteria composition and a certain disease among samples. We also provide a simple and efficient learning procedure for the proposed model using an annealing EM algorithm. Conclusion: We illustrate the performance of the proposed method both through both the simulation and real data analysis. BERMUDA is implemented with R and is available from GitHub ( https://github.com/abikoushi/Bermuda ).
Author supplied keywords
Cite
CITATION STYLE
Abe, K., Hirayama, M., Ohno, K., & Shimamura, T. (2018). A latent allocation model for the analysis of microbial composition and disease. BMC Bioinformatics, 19. https://doi.org/10.1186/s12859-018-2530-6
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.