Abstract
Transposable Elements (TEs) are DNA subsequences that have historically copied themselves throughout a genome. Apart from constituting a large fraction of all eukaryotic genomes, TEs are a significant source of genetic variation and are directly responsible for many diseases. TEs are also one of the most difficult genomic regions to analyze. A typical approach for identifying TE insertions (TEi) involves the detection of split-reads, which requires checking if each read can be split into TE and non-TE parts. Identification of the TE part depends on a model for each distinct TE class, and these classes vary significantly both within and between species. Previous methods for detecting segregating TEis depend on template libraries and their computational cost increases with the number of templates. Here we propose a novel template-free method for identifying the split-reads containing TEi boundaries called Frontier. We leverage the pervasiveness of TE sequences to identify candidate reads that might include the boundary of an insertion. We then apply machine learning methods to further classify whether the read includes actual TE-like sequence. For each predicted TEi boundary we apply a second classifier to infer the corresponding TE type (LINE, SINE, ALU, ERV/LTR). Both classifiers achieve high precision (>.9), recall (>.8) and F1 score (>.8) when applied to real data. The resulting trained model, can detect and classify about 50 million frontier reads in less than an hour. Frontier codes are available at github https://github.com/Anwica/Frontier.
Author supplied keywords
Cite
CITATION STYLE
Kashfeen, A., & McMillan, L. (2021). Frontier: Finding the boundaries of novel transposable element insertions in genomes. In Proceedings of the 12th ACM Conference on Bioinformatics, Computational Biology, and Health Informatics, BCB 2021. Association for Computing Machinery, Inc. https://doi.org/10.1145/3459930.3469545
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.