A Parallel Block Implementation of Level-3 BLAS for MIMD Vector Processors

12Citations
Citations of this article
6Readers
Mendeley users who have this article in their library.

Abstract

We describe an implementation of Level-3 BLAS 1994 based on the use of the matrix-matrix multiplication kernel (GEMM). Blocking techniques are used to express the BLAS in terms of operations involving triangular blocks and calls to GEMM. A principal advantage of this approach is that most manufacturers provide at least an efficient serial version of GEMM so that our implementation can capture a significant percentage of the computer performance. A parameter which controls the blocking allows an efficient exploitation of the memory hierarchy of the various target computers. Furthermore, this blocked version of Level-3 BLAS is naturally parallel. We present results on the ALLIANT FX/80, the CONVEX C220, the CRAY-2, and the IBM 3090/VF. For GEMM, we always use the manufacturer-supplied versions. For the operations dealing with triangular blocks, we use assembler or tuned Fortran (using loop-unrolling) codes, depending on the efficiency of the available libraries. © 1994, ACM. All rights reserved.

Cite

CITATION STYLE

APA

Daydé, M. J., Duff, I. S., & Petitet, A. (1994). A Parallel Block Implementation of Level-3 BLAS for MIMD Vector Processors. ACM Transactions on Mathematical Software (TOMS), 20(2), 178–193. https://doi.org/10.1145/178365.174413

Register to see more suggestions

Mendeley helps you to discover research relevant for your work.

Already have an account?

Save time finding and organizing research with Mendeley

Sign up for free