Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data

10Citations
Citations of this article
27Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

Motivation: fast-protein-cluster is a fast, parallel and memory efficient package used to cluster 60 000 sets of protein models (with up to 550 000 models per set) generated by the Nutritious Rice for the World project. Results: fast-protein-cluster is an optimized and extensible toolkit that supports Root Mean Square Deviation after optimal superposition (RMSD) and Template Modeling score (TM-score) as metrics. RMSD calculations using a laptop CPU are 60times; faster than qcprot and 3 faster than current graphics processing unit (GPU) implementations. New GPU code further increases the speed of RMSD and TM-score calculations. fast-protein-cluster provides novel k-means and hierarchical clustering methods that are up to 250×and 2000times; faster, respectively, than Clusco, and identify significantly more accurate models than Spicker and Clusco. © 2014 The Author. Published by Oxford University Press. All rights reserved.

Cite

CITATION STYLE

APA

Hung, L. H., & Samudrala, R. (2014). Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data. Bioinformatics, 30(12), 1774–1776. https://doi.org/10.1093/bioinformatics/btu098

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free