Ten common issues with reference sequence databases and how to mitigate them

6Citations
Citations of this article
21Readers
Mendeley users who have this article in their library.

Abstract

Metagenomic sequencing has revolutionized our understanding of microbiology. While metagenomic tools and approaches have been extensively evaluated and benchmarked, far less attention has been given to the reference sequence database used in metagenomic classification. Issues with reference sequence databases are pervasive. Database contamination is the most recognized issue in the literature; however, it remains relatively unmitigated in most analyses. Other common issues with reference sequence databases include taxonomic errors, inappropriate inclusion and exclusion criteria, and sequence content errors. This review covers ten common issues with reference sequence databases and the potential downstream consequences of these issues. Mitigation measures are discussed for each issue, including bioinformatic tools and database curation strategies. Together, these strategies present a path towards more accurate, reproducible and translatable metagenomic sequencing.

Cite

CITATION STYLE

APA

Chorlton, S. D. (2024). Ten common issues with reference sequence databases and how to mitigate them. Frontiers in Bioinformatics. Frontiers Media SA. https://doi.org/10.3389/fbinf.2024.1278228

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free