Abstract
The performance of the barrier operation can be crucial for many parallel codes. Especially distributed shared memory systems have to synchronize frequently to ensure the proper ordering of memory accesses. The barrier operation is often performed by issuing point-to-point messages and the best algorithm scales with O(logzP • L) in the LogP model. We propose a cheap hardware extension which is able to perform the task of synchronization in nearly constant time and implement a driver inside the Open MPI framework to speedup the MPI_Barrier() call. We test our implementation with the parallel version of Abinit resulting in an MPI overhead decrease of nearly 32%.
Cite
CITATION STYLE
Hoefler, T., Mehlan, T., Mietke, F., & Rehm, W. (2006). Adding Low-Cost Hardware Barrier Support to Small Commodity Clusters. In Lecture Notes in Informatics (LNI), Proceedings - Series of the Gesellschaft fur Informatik (GI) (Vol. P-81, pp. 343–350). Gesellschaft fur Informatik (GI).
Register to see more suggestions
Mendeley helps you to discover research relevant for your work.