RAP: A new computer program for de novo identification of repeated sequences in whole genomes

38Citations
Citations of this article
77Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Motivation: DNA repeats are a common feature of most genomic sequences. Their de novo identification is still difficult despite being a crucial step in genomic analysis and oligonucleotides design. Several efficient algorithms based on word counting are available, but too short words decrease specificity while long words decrease sensitivity, particularly in degenerated repeats. Results: The Repeat Analysis Program (RAP) is based on a new word-counting algorithm optimized for high resolution repeat identification using gapped words. Many different overlapping gapped words can be counted at the same genomic position, thus producing a better signal than the single ungapped word. This results in better specificity both in terms of low-frequency detection, being able to identify sequences repeated only once, and highly divergent detection, producing a generally high score in most intron sequences. © The Author 2004. Published by Oxford University Press. All rights reserved.

Cite

CITATION STYLE

APA

Campagna, D., Romualdi, C., Vitulo, N., Del Favero, M., Lexa, M., Cannata, N., & Valle, G. (2005). RAP: A new computer program for de novo identification of repeated sequences in whole genomes. Bioinformatics, 21(5), 582–588. https://doi.org/10.1093/bioinformatics/bti039

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free