CloudForest: A scalable and efficient random forest implementation for biological data

3Citations
Citations of this article
20Readers
Mendeley users who have this article in their library.

Abstract

Random Forest has become a standard data analysis tool in computational biology. However, extensions to existing implementations are often necessary to handle the complexity of biological datasets and their associated research questions. The growing size of these datasets requires high performance implementations. We describe CloudForest, a Random Forest package written in Go, which is particularly well suited for large, heterogeneous, genetic and biomedical datasets. CloudForest includes several extensions, such as dealing with unbalanced classes and missing values. Its flexible design enables users to easily implement additional extensions. CloudForest achieves fast running times by effective use of the CPU cache, optimizing for different classes of features and efficiently multi-threading.

Cite

CITATION STYLE

APA

Bressler, R., Kreisberg, R. B., Bernard, B., Niederhuber, J. E., Vockley, J. G., Shmulevich, I., & Knijnenburg, T. A. (2015). CloudForest: A scalable and efficient random forest implementation for biological data. PLoS ONE, 10(12). https://doi.org/10.1371/journal.pone.0144820

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free