Machine learning algorithm acceleration using hybrid (CPU-MPP) MapReduce clusters

Sergio Herrero-Lopez; John R. Williams

Book Chapter

Machine learning algorithm acceleration using hybrid (CPU-MPP) MapReduce clusters

Springer New York, (2014), 129-153

DOI: 10.1007/978-1-4614-9242-9_5

0Citations

1Readers

Get full text

Abstract

The uninterrupted growth of information repositories has progressively led data-intensive applications, such as MapReduce-based systems, MapReduce to the mainstream. The MapReduce paradigm has frequently proven to be a simple yet flexible and scalable technique to distribute algorithms across thousands of nodes and petabytes of information. Under these circumstances, classic data mining algorithms have been adapted to this model, in order to run in production environments. Unfortunately, the high latency nature of this architecture has relegated the applicability of these algorithms to batch-processing scenarios. In spite of this shortcoming, the emergence of massively threaded shared-memory multiprocessors, such as Graphics Processing Units (GPU), on the commodity computing market has enabled these algorithms to be executed orders of magnitude faster, while keeping the same MapReduce-based model. In this chapter, we propose the integration of massively threaded shared-memory multiprocessors into MapReduce-based clusters, creating a unified heterogeneous architecture that enables executing Map and Reduce operators on thousands of threads across multiple GPU devices and nodes, while maintaining the built-in reliability of the baseline system. For this purpose, we created a programming model that facilitates the collaboration of multiple CPU cores and multiple GPU devices towards the resolution of a data intensive problem. In order to prove the potential of this hybrid system, we take a popular NP-hard supervised learning algorithm, the Support Vector Machine (SVM), and show that a 36 ×-192× speedup can be achieved on large datasets without changing the model or leaving the commodity hardware paradigm.

Cite

CITATION STYLE

APA

Herrero-Lopez, S., & Williams, J. R. (2014). Machine learning algorithm acceleration using hybrid (CPU-MPP) MapReduce clusters. In Large-Scale Data Analytics (Vol. 9781461492429, pp. 129–153). Springer New York. https://doi.org/10.1007/978-1-4614-9242-9_5

Machine learning algorithm acceleration using hybrid (CPU-MPP) MapReduce clusters

Abstract

Cite

Register to see more suggestions