Hierarchical clustering of DNA k-mer counts in RNAseq fastq files identifies sample heterogeneities

5Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

We apply hierarchical clustering (HC) of DNA k-mer counts on multiple Fastq files. The tree structures produced by HC may reflect experimental groups and thereby indicate experimental effects, but clustering of preparation groups indicates the presence of batch effects. Hence, HC of DNA k-mer counts may serve as a diagnostic device. In order to provide a simple applicable tool we implemented sequential analysis of Fastq reads with low memory usage in an R package (seqTools) available on Bioconductor. The approach is validated by analysis of Fastq file batches containing RNAseq data. Analysis of three Fastq batches downloaded from ArrayExpress indicated experimental effects. Analysis of RNAseq data from two cell types (dermal fibroblasts and Jurkat cells) sequenced in our facility indicate presence of batch effects. The observed batch effects were also present in reads mapped to the human genome and also in reads filtered for high quality (Phred > 30). We propose, that hierarchical clustering of DNA k-mer counts provides an unspecific diagnostic tool for RNAseq experiments. Further exploration is required once samples are identified as outliers in HC derived trees.

Author supplied keywords

Cite

CITATION STYLE

APA

Kaisers, W., Schwender, H., & Schaal, H. (2018). Hierarchical clustering of DNA k-mer counts in RNAseq fastq files identifies sample heterogeneities. International Journal of Molecular Sciences, 19(11). https://doi.org/10.3390/ijms19113687

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free