Abstract
Highly accurate protein structure predictors have generated hundreds of millions of protein structures; these pose a challenge in terms of storage and processing. Here, we present Foldcomp, a novel lossy structure compression algorithm, and indexing system to address this challenge. By using a combination of internal and Cartesian coordinates and a bi-directional NeRF-based strategy, Foldcomp improves the compression ratio by a factor of three compared to the next best method. Its reconstruction error of 0.08 Å is comparable to the best lossy compressor. It is five times faster than the next fastest compressor and competes with the fastest decompressors. With its multi-threading implementation and a Python interface that allows for easy database downloads and efficient querying of protein structures by accession, Foldcomp is a powerful tool for managing and analysing large collections of protein structures.
Cite
CITATION STYLE
Kim, H., Mirdita, M., & Steinegger, M. (2023). Foldcomp: a library and format for compressing and indexing large protein structure sets. Bioinformatics, 39(4). https://doi.org/10.1093/bioinformatics/btad153
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.