Scalability of self-organizing maps on a GPU cluster using OpenCL and CUDA

27Citations
Citations of this article
28Readers
Mendeley users who have this article in their library.

This article is free to access.

Abstract

We evaluate a novel implementation of a Self-Organizing Map (SOM) on a Graphics Processing Unit (GPU) cluster. Using various combinations of OpenCL, CUDA, and two different graphics cards, we demonstrate the scalability of the SOM implementation on one to eight GPUs. Results indicate that while the algorithm scales well with the number of training samples and the map size, the benefits from using the data-parallel approaches offered by the GPU are severely limited when combined with the Message Passing Interface (MPI) in this setting, and comparable to speedups of GPU-based implementations as compared to optimized sequential code. Speedups achieved range from 3 to 32, for various map and training data sizes. We also observed a performance penalty for the OpenCL implementation as compared to CUDA. © IOP Publishing Ltd.

Cite

CITATION STYLE

APA

McConnell, S., Sturgeon, R., Henry, G., Mayne, A., & Hurley, R. (2012). Scalability of self-organizing maps on a GPU cluster using OpenCL and CUDA. In Journal of Physics: Conference Series (Vol. 341). Institute of Physics Publishing. https://doi.org/10.1088/1742-6596/341/1/012018

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free