Abstract
Motivation: fast-protein-cluster is a fast, parallel and memory efficient package used to cluster 60 000 sets of protein models (with up to 550 000 models per set) generated by the Nutritious Rice for the World project. Results: fast-protein-cluster is an optimized and extensible toolkit that supports Root Mean Square Deviation after optimal superposition (RMSD) and Template Modeling score (TM-score) as metrics. RMSD calculations using a laptop CPU are 60times; faster than qcprot and 3 faster than current graphics processing unit (GPU) implementations. New GPU code further increases the speed of RMSD and TM-score calculations. fast-protein-cluster provides novel k-means and hierarchical clustering methods that are up to 250×and 2000times; faster, respectively, than Clusco, and identify significantly more accurate models than Spicker and Clusco. © 2014 The Author. Published by Oxford University Press. All rights reserved.
Cite
CITATION STYLE
Hung, L. H., & Samudrala, R. (2014). Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data. Bioinformatics, 30(12), 1774–1776. https://doi.org/10.1093/bioinformatics/btu098
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.