kMermaid: Ultrafast metagenomic read assignment to protein clusters by hashing of amino acid k-mer frequencies

0Citations
Citations of this article
1Readers
Mendeley users who have this article in their library.

Abstract

Shotgun metagenomic sequencing can determine both the taxonomic and functional content of microbiomes. However, functional classification for metagenomic reads remains highly challenging as protein mapping tools require substantial computational resources and yield ambiguous classifications when short reads map to homologous proteins originating from different bacteria. Here we introduce kMermaid for the purpose of uniquely mapping bacterial short reads to taxa-agnostic clusters of homologous proteins, which can then be used for downstream analysis tasks such as read quantification and pathway or global functional analysis. Using a nested hash map containing amino acid k-mer profiles as a model for protein assignment, kMermaid achieves the sensitivity of popular existing protein mapping tools while remaining highly resource efficient. We evaluate kMermaid on simulated data and data from human fecal samples as well as demonstrate the utility of kMermaid for classifying reads originating from new, unseen proteins. kMermaid allows for highly accurate, unambiguous and ultrafast metagenomic read assignment into protein clusters, with a fixed memory usage, and can easily be employed on a typical computer.

Cite

CITATION STYLE

APA

Lucas, A., Schäffer, D. E., Wickramasinghe, J., & Auslander, N. (2025). kMermaid: Ultrafast metagenomic read assignment to protein clusters by hashing of amino acid k-mer frequencies. PLOS Computational Biology, 21(9). https://doi.org/10.1371/journal.pcbi.1013470

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free