Mixture models reveal multiple positional bias types in RNA-Seq data and lead to accurate transcript concentration estimates

17Citations
Citations of this article
85Readers
Mendeley users who have this article in their library.

Abstract

Accuracy of transcript quantification with RNA-Seq is negatively affected by positional fragment bias. This article introduces Mix2(rd. “mixquare”), a transcript quantification method which uses a mixture of probability distributions to model and thereby neutralize the effects of positional fragment bias. The parameters of Mix2are trained by Expectation Maximization resulting in simultaneous transcript abundance and bias estimates. We compare Mix2to Cufflinks, RSEM, eXpress and PennSeq; state-of-the-art quantification methods implementing some form of bias correction. On four synthetic biases we show that the accuracy of Mix2overall exceeds the accuracy of the other methods and that its bias estimates converge to the correct solution. We further evaluate Mix2on real RNA-Seq data from the Microarray and Sequencing Quality Control (MAQC, SEQC) Consortia. On MAQC data, Mix2achieves improved correlation to qPCR measurements with a relative increase in R2between 4% and 50%. Mix2also yields repeatable concentration estimates across technical replicates with a relative increase in R2between 8% and 47% and reduced standard deviation across the full concentration range. We further observe more accurate detection of differential expression with a relative increase in true positives between 74% and 378% for 5% false positives. In addition, Mix2reveals 5 dominant biases in MAQC data deviating from the common assumption of a uniform fragment distribution. On SEQC data, Mix2yields higher consistency between measured and predicted concentration ratios. A relative error of 20% or less is obtained for 51% of transcripts by Mix2, 40% of transcripts by Cufflinks and RSEM and 30% by eXpress. Titration order consistency is correct for 47% of transcripts for Mix2, 41% for Cufflinks and RSEM and 34% for eXpress. We, further, observe improved repeatability across laboratory sites with a relative increase in R2between 8% and 44% and reduced standard deviation.

Cite

CITATION STYLE

APA

Tuerk, A., Wiktorin, G., & Güler, S. (2017). Mixture models reveal multiple positional bias types in RNA-Seq data and lead to accurate transcript concentration estimates. PLoS Computational Biology, 13(5). https://doi.org/10.1371/journal.pcbi.1005515

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free