Complexity estimation of genetic sequences using information-theoretic and frequency analysis methods

6Citations
Citations of this article
15Readers
Mendeley users who have this article in their library.

Abstract

The genetic information in cells is stored in DNA sequences, represented by a string of four letters, each corresponding to a definite type of nucleotides. Genomic DNA sequences are very abundant in periodic patterns, which play important biological roles. The complexity of genetic sequences can be estimated using the information-theoretic methods. Low complexity regions are of particular interest to genome researchers, because they indicate to sequence repeats and patterns. In this paper, the complexity of genetic sequences is estimated using Shannon entropy, Rényi entropy and relative Kolmogorov complexity. The structural complexity based on periodicities is analyzed using the autocorrelation function and time delayed mutual information. As a case study, we analyze human 22nd chromosome and identify 3 and 49 bp periodicities. © 2010 Institute of Mathematics and Informatics.

Cite

CITATION STYLE

APA

Damaševičius, R. (2010). Complexity estimation of genetic sequences using information-theoretic and frequency analysis methods. Informatica, 21(1), 13–30. https://doi.org/10.15388/informatica.2010.270

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free