Fast parallel algorithms for blocked dense matrix multiplication on shared memory architectures

G. Nimako; E. J. Otoo; D. Ohene-Kwofie

Conference Proceedings

Fast parallel algorithms for blocked dense matrix multiplication on shared memory architectures

Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (2012) 7439 LNCS(PART 1) 443-457

DOI: 10.1007/978-3-642-33078-0_32

4Citations

7Readers

Get full text

Abstract

The current trend of multicore and Symmetric Multi-Processor (SMP), architectures underscores the need for parallelism in most scientific computations. Matrix-matrix multiplication is one of the fundamental computations in many algorithms for scientific and numerical analysis. Although a number of different algorithms (such as Cannon, PUMMA, SUMMA etc), have been proposed for the implementation of matrix-matrix multiplication on distributed memory architectures, matrix-matrix algorithms for multicore and SMP architectures have not been extensively studied. We present two types of algorithms, based largely on blocked dense matrices, for parallel matrix-matrix multiplication on shared memory systems. The first algorithm is based on blocked matrices whiles the second algorithm uses blocked matrices with the MapReduce framework in shared memory. Our experimental results show that, our blocked dense matrix approach outperforms the known existing implementations by up to 50% whiles our MapReduce blocked matrix-matrix algorithm outperforms the existing matrix-matrix multiplication algorithm of the Phoenix shared memory MapReduce approach, by about 40%. © 2012 Springer-Verlag.

Author supplied keywords

Cite

CITATION STYLE

APA

Nimako, G., Otoo, E. J., & Ohene-Kwofie, D. (2012). Fast parallel algorithms for blocked dense matrix multiplication on shared memory architectures. In Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics) (Vol. 7439 LNCS, pp. 443–457). https://doi.org/10.1007/978-3-642-33078-0_32

Fast parallel algorithms for blocked dense matrix multiplication on shared memory architectures

Abstract

Author supplied keywords

Cite

Register to see more suggestions